Now for live streaming, all you need is one model!?
You didn't mishear, yesterday a live streaming room that used AI to generate content in real-time actually emerged.

This isn't the kind of "semi-AI" digital human livestream, as this livestream room doesn't have a host or a script, with content determined by audience requests in real-time.
And it's not pre-made food—after you place the order, AI prepares it fresh on the spot.
The model used in this live broadcast can generate a 5-second video in under 3 seconds, with the next segment ready before the previous one has even finished playing.
So, although the dishes are prepared on the spot, the "serving" process of this live broadcast has been non-stop from the start of the show until now.
In addition to live streaming, developers have also used Codex to create an "AI version of Douyin", where the content is generated in real-time by video models and seamlessly connected.
You're still browsing the video on the screen, and the next segment has already been generated and is ready to switch to seamlessly as you swipe over.
If a short clip isn't enough, the video can be extended to 30 seconds for a truly striking effect.

The model powering this real-time livestream and AI-driven short-video app is the H3 Max, which MiniMax just launched.
Just click whatever you want to see, and the AI will generate it for you on the spot.
H3 Max comes from a joint development by MiniMax and Fal.
MiniMax open-sourced its video generation model H3 at the end of July, which garnered 24 million downloads in just three weeks.
Fal is a company focused on AI inference infrastructure, specializing in inference acceleration, with extensive experience in making models run faster.
The two parties joined forces and collaborated on a project.
Put simply, the work builds on H3, adding post-training and inference optimization tailored specifically to real-time generation scenarios.
In simple terms, it means equipping the H3 with an acceleration motor as well.
The resulting product is called H3 Max - the core change is one word, fast.
Generating a 5-second 768p video takes less than 3 seconds with the H3 Max, with a throughput approximately 35 times that of the original H3.
The generation speed has already surpassed the video playback speed, with the next segment ready before you've even finished watching the previous one.
Fast as it may be, quality hasn't slipped in the slightest.
H3 Max ranked first in video generation on the two mainstream lists, Artificial Analysis and Design Arena.

Quantitative changes accumulate and trigger qualitative changes. When the time it takes to generate a video is shorter than the video itself, a whole new way of playing emerges that previously did not exist.
The first person to eat the crab was Rehan Sheikh.
He is an engineer at Fal, able to come up with new ideas first, and also benefits from being close to the action.
He connected the H3 Max to an overseas live streaming platform and created a channel called "Cross-Dimensional TV".
The main selling point of this channel is that it doesn't have a main selling point.
The screen may be playing an animated short one second, then suddenly cut to an absurd 1960s puppet commercial the next, and switch to a documentary about outer space the second after that.
There is no program list and no pre-recorded material; all images and sounds are generated in real-time by the model, advancing the live stream segment by segment.
Afterwards, he posted clips of the live broadcast on X, garnering 5.4 million views.

Closely followed by independent developer Pieter Levels.
This individual is a celebrity in overseas developer circles, having previously created the digital nomad community Nomad List and the AI photo editing tool PhotoAI, and is skilled at single-handedly taking products from development to launch.
The on-demand AI live broadcasting website mentioned at the beginning is the brainchild of Pieter.
Almost at the same time, another engineer of FAL, Alex Koumpas, also started a live stream.
What's different is that his live broadcast tells a continuous story.
Of course, the audience can also intervene at any time by typing instructions in the chat box, such as "let the protagonist turn around and walk into the forest" or "sudden red rain starts falling from the sky".
Once the model receives the command, the visuals shift course, and the narrative follows the audience's lead, veering in a new direction.

Within a single week, three people—without consulting one another or making any prior plans—independently launched three AI livestream rooms.
The H3 Max's speed took center stage, and everything fell into place.
OpenClaw Moment in Video Generation
A faster video model generation speed is certainly a good thing, but merely increasing speed does not necessarily constitute a moat.
"I's a extremely vibrant top of a top of top of top of "I's a extremely vibrant top of top of top of "I's "I's top of top of "I only top of "I's common. "I's only top of " We only top of "I only top of "I only top in "I only top of "I only top of " "I only top of "I only top of "I " only top of "I It's only top of "I only top of "I only top of "It's only top of "It's "I only top of "I only "I only top of "I only top of "I "I only top of "I only top of "I "I only top, "I only top of "I "I " "I "I "I " "I " "I " "I " "I " "I " "I " "I " "I " "I " "I " "I " "I " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " " "
The new "AI real-time live streaming" model introduced by H3 Max could mark a watershed moment for the video generation sector.
It directly eliminates the intermediate waiting process.
This new model has directly opened up an entire path to commercialization.
Initially, when Peter Steinberger was working on OpenClaw, he transformed the AI Agent from a command-line tool that only developers could use into an open-source assistant that ordinary people could easily access.
This idea wasn't supposed to be so simple, but after OpenClaw's download volume reached 4.7 million, Sam Altman and Zuckerberg could no longer sit still.
Real-time AI video points to the same structural change, making "unattainable" technology accessible to everyone.
However, video generation and coding volumes are fundamentally different, with completely different business structures.
The code has right and wrong answers, with standard solutions, resulting in severe homogenization of code models and collective involution within the industry.
But the video does not have a unique correct answer.
Aesthetics are subjective, and when the same prompt is given to ten different models, ten different styles emerge, with users choosing based on "whose image I prefer".
This means that the video generation track will not converge to one or two winners, and competition will not rely solely on price.
Moreover, the stickiness will deepen with increasing use.
A creator uses a certain model to create their own visual style, which the audience recognizes and the brand buys into.
Once the model is switched, it means that the entire set of visual assets and audience expectations will have to change accordingly.
As usage accumulates, the model delves deeper into the material library, character settings, and creative workflows, and migration costs will continue to rise.
Open-source and open ecosystems have added another layer of barriers.
Community developers and creators build derivatives around the same base, with feedback feeding back into model iterations.
Once this model is up and running, it will be difficult for latecomers to catch up solely by relying on model capabilities.
H3 has already established the initial form of this flywheel effect.
Open-sourced for three weeks, with 24 million global downloads, over 300 derivative models, including FastH3, H3 480p, H3 Max...
H3, as a branch, has been nurtured by different teams and has extended in various directions, ultimately blossoming everywhere.
Moreover, the capital market is already voting with its money.
Last week, Morgan Stanley maintained its "overweight" rating on MiniMax, with a target price of HK$900.
Today, MiniMax's Hong Kong stocks rose 16.18%, with the stock price surging from HK$300.4 to HK$349.
Market performance never lies, and these examples are enough to reflect that the capital market is reevaluating MiniMax's commercialization capabilities.
The father of OpenClaw never thought that what initially started as a tool for his own use would, after being open-sourced, ignite a global Harness war.
MiniMax had not anticipated that previously no one had used AI to generate videos in real-time during live streams, but now someone is using the H3 Max to operate a nearly non-stop TV station.
MiniMax's decision to open-source its state-of-the-art video model amounts to a wave of "technology democratization."
They have directly placed cutting-edge capabilities in the hands of every developer, elevating the foundation of the entire track to a new level.
As global developers and creators run in their own directions on this foundation, accumulating materials, workflows, and user habits around the same ecosystem, the flywheel of commercialization will no longer need anyone to push it.