ChinaChina
CSI 3004,569.79 0.48%
Hang Seng25,387.76 0.23%
Shanghai3,957.82 0.42%
CNY/USD6.7174 0.04%
FEATURE

8 min read

Mysterious Deepfake Team Releases Explosive Demo Videos, Exposing Self-Evolution Model Tech Roadmap

Do you still remember the mysterious 10-minute single-take video that circulated on Weibo last week?

As soon as the video came out, netizens expressed: unbelievable, as if it were something from outer space...

It even received the highest praise: This must be AI-generated, right?!

Peers came to inquire about which company it was, and it's said that a leading company even held a meeting to study it... our backend was flooded with questions from netizens.

After watching this 10-minute one-take video, people are saying that the era of embodied intelligence with ChatGPT may have truly arrived, with some calling it a major milestone for embodied intelligence and others referring to it as an innovative architecture, as it has shown:

Of course, due to the robot's seemingly "sci-fi" performance, some skeptics have questioned whether it was remotely controlled. (Although the team said that in such a small space, they really didn't know how to achieve remote control.)

And this mysterious team responded to the doubts in a unique way: they released several more demos, still rough, still shot in a single take, as if to say:

"Intelligence, whether you believe it or not, is really on its way."

Thanks to the first article, we were able to get in touch with the mysterious model team and expressed our question - why not come out and claim it directly. The response we received was: the model is like a person, let it speak for itself.

"Be a decent person" large model.

And several demos released this time show that the robot has "really become human-like", thinking like a human, making decisions independently, moving smoothly, and being logically consistent. Moreover, according to the team, they will soon let the robot go out on the street, in an open environment, to be seen by everyone directly.

Video 1: Seamless Hand-Eye Coordination Demonstrates Advanced Kinematic Understanding

If you're holding a bottle of water and someone pokes your arm, can you react quickly enough to prevent the water from spilling?

A robot held a cup of water in its left hand, and when pushed by a human with a stick, it dodged to the other side and then returned to its original position, without spilling a single drop of water from the cup.

If you're holding a bottle of water and someone pokes your arm while also pushing over a can next to you, can you react quickly enough to prevent the water from spilling and the can from falling?

A human pushed three cans on the table from the right side, and its right hand simultaneously steadied the cans.

The robot performed like a tai chi master, able to neutralize its opponent's blows and apply force with precision.

Moreover, upon closer inspection, when the can is pushed, the left arm holding the water has not yet recovered, and the robot uses its right hand to support the can.

This feat has already surpassed many humans, to say nothing of robots; everyone can try to see if they can complete these two tasks simultaneously with both hands.

Mainstream VLA and WAM models have a flaw: once they encounter scenarios that require physical reasoning, such as occlusion, contact, and deformation, they tend to fail. This is because they are merely imitating the appearance of actions without understanding the underlying physical laws.

The "Be Yourself" model seems to have equipped the robot with a "physical brain". It is the unification of dynamic prediction and long-term state that allows the trajectory to naturally satisfy dynamic feasibility.

In terms of action response speed and smoothness, the robot's movements are no longer based on "guessing" through memorized data, but rather on "learning" through an understanding of physical rules, similar to humans.

It's the difference between "imitating blindly" and "understanding the underlying principles".

Video 2: Understanding Spatial and Object Attributes, and Summarizing Human Habits

In the second video, a humanoid robot formally takes over household chores. The team behind the "Be a Person" model says this is a video of the robot completing the task with zero shots.

They said the model's zero-shot success rate has reached 80%.

I named this video "Forced Cleanliness Robot: The Living Room Must Not Be Messy" (doge)

The robot first put the small toy it was holding back into the entrance hall's storage rack, making a "returning things to where they came from" gesture. Then, it pulled the storage basket filled with miscellaneous items to its side, slightly crouched down, and one after another, took out a black hat and a purple fitness ring, hanging them stably on the coat rack.

Finally, it pulled out a gigantic flower-shaped cushion from the basket, which was apparently flexible, but this time it simply tossed it, and the cushion landed on the sofa.

Throughout the entire process, the robot's movements may not have been particularly agile, but they gave a sense of being entirely under control.

The robot's understanding of space and object properties has become very similar to that of humans: it knows that objects can be dragged on smooth floors, ring-shaped objects like hats and exercise hoops can be hung, and flat cushions will unfold on their own.

They are even "lazy" to the point of throwing things directly instead of bending down to pick them up, proving that laziness drives scientific development and lazy robots drive intelligent evolution.

Mainstream models have yet to achieve human-like generalization capabilities. VLA relies on "brute force" to calculate, and it easily collapses when encountering unfamiliar scenarios; WAM is like building with blocks, piecing together actions in a rigid manner, and it fails to adapt when slightly modified.

The robot equipped with the "做个人吧" model appeared highly stable, integrating spatial and object-property understanding with human habits to successfully complete a long-horizon task.

Video 3: Robots Team Up to Smooth, Shake, and Place Sheets

When we were young, our elders would often call us over to give freshly washed bed sheets a good shake.

In this video, two different models of robots appeared, and it is said that all their training was based on watching human videos, so they once again performed "one brain with multiple forms" cross-entity collaboration, and actually worked together as a team.

All that can be seen is a bed sheet that appears to have just been taken out of a dryer, with one end being pulled by a taller robot and the other end being pulled by a shorter robot.

The shorter robot took a step forward and bent down, pulling the sheet taut. Immediately, the taller robot gently shook the sheet a few times with both hands.

Although the robot's poses in the video are still somewhat stiff, their behavioral logic is surprisingly human-like. It's no different from when we give freshly washed clothes a few shakes to prevent wrinkles.

It's worth noting that mainstream models have a fatal flaw: they are not like humans and cannot truly achieve self-evolution. VLA can only make minor adjustments and cannot change its "circuitry"; WAM is too rigid and cannot adapt to a new environment.

In contrast, the MVP model appears to unify understanding of the physical world with understanding of human logic.

Video 4: Energy-Driven, Optimal Energy Behavior Solutions

When robots learn to think and understand that it's more efficient to pick up and put away everything at once, rather than making multiple trips back and forth, what will they do?

In the video, the robot's operation of storing items is like a human's: no matter how many items there are, it tries not to make a second trip.

It first pulled out a black scarf from the basket and draped it around its neck like a shop assistant. Then it picked up a purple ring and wrapped it around its forearm, even lifting its hand to let it slide down to the elbow, successfully transforming itself into a walking clothes rack.

Then there were slippers and headphones, and after a series of maneuvers, the robot's body was laden with items, yet it calmly turned around and walked away, leaving behind a light and airy figure.

A one-minute video shows the robot picking up four items, not only fully demonstrating its ability to handle long-range tasks, but also showing a deep understanding of each item and its ability to establish a corresponding mapping with its own capabilities.

To put this into perspective, current mainstream models performing the same tasks would likely do so one by one, as they are essentially based on statistical fitting and guessing, lacking human-like understanding of the world and autonomous intentions.

If the robot in the video saw this, its internal OS might be thinking: "Come on, you've come all this way, how can you only take that little bit? Taking more would be more efficient!"

The "doing the right thing" model's robot makes decisions as if it doesn't need to deliberately "think" about what to do next, it just naturally slides in the direction of energy decrease, like water flowing downhill.

In this "downhill" process, thinking, exploration, and action are seamlessly connected.

Perhaps this is what true intelligence—one that genuinely understands the world and possesses autonomous intent—looks like.

After watching these 4 demos, I finally understood why the model is called "be a person". The robot's every move indeed resembles a person's decision-making process, and in long-term tasks, interfering environments, and collaborative work, it demonstrates continuous learning and evolution regarding the physical world, human behavioral logic, and its own capabilities.

Robots can understand space and object properties like humans, and their thinking logic and operational habits when doing household chores are similar to those of humans. They can also self-drive, autonomously learn, and continuously self-evolve like humans.

Moreover, robots also embody these qualities:

Other features are too numerous to mention, but I'm now truly starting to believe that the moment when embodied intelligence can be truly utilized by humans is approaching.

Intelligence may have actually emerged in places we can't see.

What do you think of these new demos? Feel free to leave a comment in the discussion section, let's interact and provide some encouragement and pressure to prompt the tech team to explain themselves, which will also facilitate communication.