CSI 3004,552.58 0.10%
Hang Seng25,213.31 0.46%
Shanghai3,942.09 0.02%
CNY/USD6.7088 0.17%
FEATURE

10 min read

World Generation Models Arrive for 3D Leaders: Scene-Level Generation Enters Production Pipelines

NIU has taken off. 3D generation has just taken a big step from "creating things" to "creating worlds".

Have you seen "Inception"? The dream weaver can fold streets under the laws of physics with a wave of his hand, and buildings rise accordingly.

A similar scene is now unfolding in 3D generation:

Upload any scene image, and a complete 3D scene will be generated directly within 2 to 3 minutes.

The bowls on the table can be lifted individually, the chairs can be rotated and moved, and unwanted furniture can be replaced at any time and anywhere...

Unlike three-dimensional images that can only be observed, every asset here is "alive".

Alternatively, in addition to niú lái, we can also have mǎ lái or hǔ lái.

The company taking this step is Hyper3D, a leading player in 3D generation, which has officially launched WorldGen, a world-generation model that pushes 3D generation from individual assets to complete scenes.

In comparison to the broad "world model" concept prevalent in the industry, WorldGen's positioning is more specific and clear-cut —

It's not about creating a world that can only be viewed, but rather delivering 3D scenes that are interactive, editable, and can be entered into real production environments.

It's like building with Lego blocks, where you can continuously add modules inside and remove or replace any block at any time; each block also has its own physical properties, and if you accidentally remove a crucial supporting object, the surrounding objects will also shift or fall.

In simple terms, each asset here does not completely lock with the background, and the physical relationship between assets is also clear.

It is also more practical and closer to real work.

Games need levels, films need sets, and robots need training environments... but now they can all start with WorldGen, whose generation capabilities are already sufficient to be used in real production scenarios.

This road didn't come out of thin air, as early as last year, Yingwu Hyper3D won the best paper award at SIGGRAPH 2025, the top international graphics conference, with its self-developed scene-level generation technology CAST.

The only other business companies to receive awards at the same time were Google and Meta.

While CAST is the technological foundation of WorldGen.

So if we say the first step of AI 3D is to answer what an object looks like, then WorldGen answers the second question:

Multiple objects, each with their own spatial relationships and physical attributes, can be combined to form a scene.

From One Image to Multiple Assets and a Full Scene

However, most current overall scene generation methods have a common problem:

Looks good, but is not very useful.

Users can do very little within the scene, basically just walking around and looking. Want to make changes? Sorry, you'll have to start over from scratch.

The problem lies in the generation method, according to Wu Di, founder and CEO of Yingyu Hyper3D:

Traditional methods directly generate the entire scene as a single model, making it easy for all content to be merged into an indivisible whole, and the scene is difficult to reuse.

The key change for WorldGen is to break down scenes into individually operable assets while preserving their spatial and physical relationships.

The entire generation process can be broken down as follows:

Open the official website (https://hyper3d.ai) and select WorldGen to generate.

Here we upload a cartoon-style partial kitchen diagram.

The system can automatically identify the main objects in the image, and also allows manual selection to specify the parts to be generated.

Next, each foreground object will be generated as a separate asset, with a preview available in approximately 2 to 3 seconds.

The background environment is filled in using 3D Gaussian Splatting to preserve visual completeness. The overall piece can be compressed to two to three minutes.

When SimReady mode is enabled, the system also adds collision bodies to each asset and estimates the physical properties required for simulation, including mass, friction coefficient, and restitution coefficient.

They have finally successfully replicated the 3D version of Huo Nao Chufang.

With the initial draft of the scenario in hand, industry applications are also within reach.

First, let's look at the field with the most urgent demand for data, namely embodied intelligence.

Real-world data collection not only incurs costs related to venues, equipment, and personnel, but also cannot safely replicate hazardous environments such as fires or collapses. Traditional simulation methods, while able to fill the gap, require engineers to spend a significant amount of time manually setting up scenarios, resulting in high costs and large deviations.

WorldGen has achieved a subtle balance between the two.

Take the closed-loop solution jointly released by Yingmou Hyper3D, Digua Robotics, and Mou Xianfei in July this year as an example:

WorldGen is responsible for generating interactive 3D scenes from real images, with all 3D objects in the scenes having simulated import capabilities.

Mogu's MotrixSim engine is responsible for physics and parallel simulation, allowing robots to move, grasp and interact within it.

Digua Robot provides computing power and development toolchains.

It is only through the integration of three parties that a complete chain can be formed, ranging from scenario preparation and simulation operation to data collection and strategy verification.

As a member of NVIDIA Inception, Yingwu Hyper3D is also further connecting with NVIDIA's global simulation ecosystem through WorldGen, expanding "single asset for Isaac Sim" to "entire scene for Isaac Sim".

The gaming circle is more straightforward in its approach.

Obtaining a level concept map is just the beginning, and the art and planning teams still have a long string of work ahead of them.

Tusong 3D can already shorten the production time of individual assets, but these assets still need to be assembled one by one by humans.

WorldGen is attempting to push this part of the work further upstream.

It starts by generating an objectified scene from the concept art, allowing planners to review the movement flow and spatial layout, level designers to move or replace props, and the art team to decide which key assets merit detailed refinement.

What's more convenient is that the standalone assets generated by WorldGen can directly enter DCC and real-time engine environments such as Blender, Unity, PlayCanvas, Unity Engine, and Unreal Engine, for secondary creation including material adjustment, animation binding, level editing, and performance optimization.

In July this year, Yingmou Hyper3D has partnered with Unity China to jointly build an engineering closed loop that connects 3D generation to real-time game applications.

According to the team, during WorldGen's gray testing phase, many game industry practitioners have also expressed interest in it.

The film and television industry sees its potential as a "control room".

Wu Di gave a very specific example:

The director wants to move the apple from the plate on the left side of the screen to the right side, which is difficult to control stably with pure video generation.

Using WorldGen to build a controllable 3D foundation and then enhancing the final image quality with a video generation model is a more direct approach.

Some creators have already combined WorldGen with Seedance 2.5:

WorldGen provides 3D scenes, allowing for easy camera scheduling and composition adjustments, before handing them over to video generation models to fill in character performances, material lighting and shading, and final styles.

WorldGen's AI Render module features AI rendering and free camera movement, allowing the same scene to be transformed into multiple visual styles, thereby bringing a dual boost to production efficiency and quality for content such as product concept demonstrations and space tours.

This new application will enable video generation to achieve both controllability and visual appeal, producing blockbuster-level results from the outset.

It is understood that WorldGen has now entered the production process for actual film and television projects, with related works to be launched online after completion.

XR and spatial computing have verified its versatility.

Since scenes can be observed from any angle and main objects can be replaced and moved, an ordinary photo has the opportunity to become an immersive space that can be entered, browsed, and adjusted.

Indoor design, immersive education, and spatial display can all be its potential applications.

Paired with XR devices like Apple Vision Pro, "Ready Player One" is coming into reality.

Of course, these are just a small part of the uses of WorldGen, with infinite potential waiting for you to explore.

From Best Paper to Product: WorldGen's Year-Long Journey

To understand how WorldGen rose to prominence in one step, it's necessary to revisit the CAST paper from last year.

Scene generation is difficult because it requires the model to have a deep understanding of the context, semantics, and spatial relationships between objects, as well as the ability to generate coherent and realistic scenes. This involves complex tasks such as object recognition, scene layout prediction, and image synthesis, which are challenging to achieve with high accuracy and consistency.

It's not something that can be achieved by simply running a 3D rendering continuously for over a dozen times. A single image inherently lacks depth, true scale, and information about the back of objects, and is also subject to occlusion and perspective distortion.

For instance, if an apple is occluded by a plate, the model must first complete the invisible parts; if it generates a slightly larger object, it may pass through the plate when placed back on the table; if the position is slightly offset, subsequent collision bodies and support relationships will also go wrong.

Incorrect scenarios can also cascade, with one mistake leading to another, ultimately resulting in real costs for users to rework products.

The CAST approach can be condensed into four steps: first, understand; then, fill in the gaps; next, put it back; and finally, perform physical error correction.

Step 1: The model breaks down the input image into objects and estimates relative depth, transforming an unstructured RGB image into an object list and spatial cues.

Step 2: For each object, independently generate a complete 3D geometry. The occlusion perception mechanism will refer to the visible parts and scene information to fill in the areas that are not visible in the image.

Step 3: Place the generated objects back into a unified space. The system needs to estimate rotation, translation, and scale, and try to maintain consistency with the original image's composition and positioning.

Step 4: Build a finer-grained object relationship graph to describe contact, support, and suspension relationships, then apply physical correction using SDF to reduce interpenetration, floating objects, and illogical stacking.

This framework won the best paper award at SIGGRAPH last year, with the organizing committee commenting:

CAST supports open-vocabulary reconstruction tasks, excelling in handling occlusions, precisely aligning objects, and ensuring consistency between the physical world and input images, opening up new possibilities for virtual content creation and embodied intelligence.

But CAST still has a ways to go before it can be productized.

The team later spent nearly a year specifically addressing the issue of error accumulation caused by the serial connection of multiple modules.

By reducing redundant modules, restructuring underlying processes, and enhancing the stability of individual links, WorldGen was finally released in its complete form.

We believe that world generation models can recreate scenes depicted in text prompts or images as usable 3D worlds.

It can generate open-ended scenes, and it is also closer to the expression of a "world."

It remains a key component of the generalized world model — after all, a true "world model" needs some medium to carry the world's appearance and attributes. Yingmou simply chose 3D as that medium.

3D Generation Enters Scene-Level Era

Returning to the original question: What exactly did WorldGen solve?

The answer is already clear in the application.

It won't directly generate a game that can be launched online, a complete movie, or complete the training and decision-making for embodied intelligence, but it solves the common underlying difficulties faced by all these industries:

How to combine creative ideas with massive assets and quickly monetize them through scenario-based applications.

3D generation models, led by Hyper3D Rodin Gen-2.5, have cut the production cycle for individual assets from days to minutes or even seconds. Yet in real-world projects, standalone assets are rarely all that's needed.

Subsequent steps will inevitably involve assembling assets of different forms and physical properties, including identification, calibration, and importing into the engine, all of which still require human intervention.

What WorldGen changes is precisely the fundamental delivery unit of AI 3D:

From a generated object to a roughly constructed scene.

While the final product still has room for improvement, it already provides an editable starting point for game designers, film creators, and simulation engineers.

But for the scenario to be truly usable, merely "looking like" something is not enough.

The simulation of physical laws determines the upper limit of WorldGen, and the more detailed the creation, the more realistic and useful it is.

When the generated scenes can carry content, simulate interactions, and access existing workflows, they attract not only ordinary users who try new things, but also professional creators who truly need to improve production efficiency.

3D generation is one of the AI applications closest to the production process. Its value will only be truly unleashed when the generated results can be continuously edited, reused, and put into production.

What used to limit the application of 3D generation is the scene-level gap that WorldGen is now filling.

Its emergence can be said to have taken advantage of a perfect combination of timing, location, and harmony among people.

Thus, from generating objects to assembling worlds, 3D generation has entered the era of scene-level creation.

The official website link is: https://hyper3d.ai