CSI 3004,552.58 0.10%
Hang Seng25,213.31 0.46%
Shanghai3,942.09 0.02%
CNY/USD6.7088 0.17%
FEATURE

9 min read

Chen Dawei Returns, Enters the Large Model Sector

Something big has happened - a dark horse has emerged among China's top domestic large model developers.

A new company, with its first model having only 27B parameters, took second place overall in the China Academy of Information and Communications Technology's MCP specialized test.

What came before it was the ace that Liang Wenfeng had been hiding for nearly a year - the DeepSeek-V4-Pro with a parameter quantity of up to 160 billion.

The difference between the two is a mere 1.3 percentage points.

This dark horse is StartLux-V1.0-27B-Preview, from StartLux (formerly Yuandian Xinghui).

Looks a bit unfamiliar, doesn't it? Don't worry, its founder and CEO is an old acquaintance: Chen Danyan.

Known as the "godfather" of programmers in the internet era, his achievements are a matter of public record:

Shanda Network's co-founder and the head of Lian Shang Network, one of China's earliest programmers to introduce the concept of "shareware"

After a decade of retirement, he has made a comeback, this time placing his bets on local models.

This has to do with Chen Dawei's recent rare public appearance.

He publicly spoke out at the 18th anniversary reunion of Shanda Innovation Institute, saying "there are eight non-consensus points in the AI era", with four of them emphasizing local models.

Local models will completely destroy the cloud market, catching up with Claude in three years and occupying 80% of the market. The model competition based on parameters is about to become outdated...

StartLux is the best representation of his idea.

The company's business has not followed the industry trend, instead focusing on the commercialization of small and beautiful local models, making it the world's first truly local model company in the true sense of the word.

As the first market-oriented scorecard, StartLux-V1.0-27B-Preview does not rely on the cloud and can run directly on consumer-grade PCs.

In other words, this local model, which is nearly 60 times smaller, has Agent capabilities that can match those of the trillion-scale cloud flagship.

What's the basis for this?

Small Model Achieves Big Results, 27B Outperforms 1.6T

Before the answer is revealed, let's take a look at who the comparison is being made to.

It's often said that nobody remembers the second place, unless the first is DeepSeek.

Moreover, the gap is only a hair's breadth, making it well worth discussing.

The results come from the authoritative institution, China Academy of Information and Communications Technology's trusted AI large model benchmark test MCP special item, which sets six types of tasks around real application scenarios:

Location navigation, web search, browser automation, financial analysis, code repository management, 3D design.

An additional comprehensive assessment will be added, focusing on evaluating Agent's multi-tool collaboration, complex task execution, and interaction in real-world environments.

In simple terms, MCP-Universe doesn't care about how well a model responds, but rather whether the task can actually be accomplished. This is also the most essential aspect in evaluating the quality of an Agent.

The results showed that StartLux-V1.0-27B-Preview had a comprehensive score of 39.25%, ranking second.

With over 284 billion parameters, DeepSeek-V4-Flash-0731 and 198 billion parameters for Step-3.7-Flash, it trails DeepSeek-V4-Pro by just 1.3 percentage points.

With the same 27B parameter scale, StartLux also surpasses Qwen 3.6 by 5.34 percentage points.

It also excels in individual subjects, ranking first in location navigation, financial analysis, and browser automation, with other sub-items also ranking high.

Let's take a look at two case studies, putting data aside.

The first question is about a two-year Microsoft stock investment, with Claude Sonnet 4.6 as the competing topic.

Claude's answer is: $47,254, 89.02%.

Video link: https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfg

StartLux gave: $47,499.09, 90.00%.

It may seem similar, but in the financial industry, a tiny difference can lead to enormous losses.

A closer look at the thinking process of the two models reveals that due to missing original data, Claude Sonnet 4.6 directly misjudged January 8, 2025, as a non-trading day and calculated the closing price of the previous day instead.

Under the same circumstances, StartLux retrospectively reviews the original data and verifies the market trends around the target date, confirming the accurate closing price before completing the calculation and generating visualization.

The conclusion ultimately drawn by StartLux is the correct one, and is verifiable and traceable.

The second question is more intuitive, having two models assist simultaneously in querying flight tickets in a browser.

Open the browser and go to Google Flights, find a one-way flight from Singapore to Beijing departing 5 days later, I want the cheapest direct economy class flight that does not land at Daxing Airport, only consider the price, and close the browser after completing the task.

Among them, StartLux-V1.0-27B-Preview took about 95 seconds to complete the search and found a China Airlines ticket priced at $299.

In contrast, Claude Sonnet 4.6 experienced more screenshots confirming issues with date selection, pop-up window closure, and filter menus, and at one point accidentally touched the time filter panel.

More than 200 seconds later, it gave a lowest quote of $556. The price was higher, and the search time was still twice that of StartLux.

Especially in terms of operational pathways, StartLux is much more concise, requiring only 12 steps, whereas Sonnet requires a full 21 steps.

This is enough to show that on the Agent task, the scale of parameters is no longer the only decisive variable.

A new variable is post-training.

Cutting Prices, Not Capabilities

StartLux-V1.0-27B-Preview was not trained from scratch.

It is also based on Qwen3.6-27B, but the final test score is significantly higher than Qwen, and the reason lies in the task data and automated post-training methods.

In simple terms, StartLux is training a more capable Agent, with a focus on reinforcing its understanding of tasks, tool selection, parameter construction, multi-step execution, status checking, and result verification.

The model needs to learn not only to output the final text, but also when to call which tool, how to adjust when the tool returns an exception, and under what circumstances it can declare the task complete.

Further training will push this process even further.

The team has independently developed a brand-new, multi-dimensional, verifiable, and scalable model iteration optimization technology, which uses the AI-trained AI (Auto Research) method, allowing the model to autonomously execute tasks in a real-world tool environment and continuously adjust its strategy based on environmental feedback.

For instance, the financial analysis case mentioned earlier, which involves backtesting and revision, as well as the constraint identification and path selection in browser tasks, are the most intuitive manifestations of post-training.

According to official information, StartLux-V1.0-27B-Preview is also the country's first local Agent model to complete post-training using the Auto Research method.

This does not mean that the Scaling Law is invalid.

Large parameter cloud models are still the mainstream choice at present, but StartLux has also given a clear signal: this is not the only solution.

In the words of Chen Danyan:

Scaling Law is a "passing fairy," without it, AI cannot take off, but it has merely passed through the development path of AI and its future is not necessarily tied to it.

The industry has also become aware of this issue, and the technical path of large models is currently showing a trend of distinct divergence:

On one hand, there are the faithful believers in "more power leads to miracles," with parameters scaling from hundreds of billions to trillions, and then to tens of trillions, while training costs also surge exponentially;

On the other hand, the Agentic evaluation system, represented by Harness, has quickly become popular, with more and more experts and scholars explicitly advocating for "less is more".

To paraphrase Wang Yangming, knowing and doing are one, and higher-quality action is the key. StartLux has provided yet another example of this by streamlining its model.

This can also explain why StartLux insists on a local model.

Once the model is on a PC, for long-term tasks, its cost structure can shift from continuously accumulating cloud-based token fees to more controllable device computing power and electricity consumption, while the model can also better understand the user's long context, achieving more personalized goals.

It's not just StartLux, as Meta, Google, and NVIDIA have also recently been increasing their investment in local small model development.

However, most of these are still in the experimental stage or cater to niche groups, with only one company focusing on local model commercialization.

StartLux also stated that they will steadily advance their own foundation model training and explore new architectures such as diffusion-based language models, and as long as they are on the right path, the future looks promising.

StartLux has taken over the local model, and since this is a non-consensus route, the people at the helm need to have two qualities: the courage to make bets and the ability to deliver.

StartLux's star team is a case in point, with entrepreneurs + scientists forming a strong alliance.

StartLux founder Chen Danyan

CEO Chen Danyan had previously been briefly introduced as a serial entrepreneur and one of China's first-generation programmers, who rose to fame around the same time as Zhang Xiaolong and Lei Jun, and started his business ventures in the same era as Ma Yun and Ma Huateng.

He was one of the first people in China to introduce the concept of "shared software" and co-founded Shanda Network with his brother Chen Tianqiao, as well as Shanda Innovation Institute, one of the cradles of internet innovation, and also incubated the nationally popular product "WiFi Master Key".

It can be said that he is extremely familiar with Chinese market users and products, and local models are also his comfort zone.

StartLux co-founder Guo Quanwei

The person responsible for implementing the technology roadmap is StartLux co-founder and CTO, Guo Quanwei.

Kuo Chuan-wei holds a Ph.D. in Computer Science and Engineering from National Yang Ming Chiao Tung University, with research areas covering locally deployed large models, Agentic AI, AI for Science, AI for Finance, and privacy-preserving machine learning.

Before joining StartLux, he served as the chief algorithm scientist at AI Science, a company called Huanliang Technology, and earlier worked at Taiwan's Industrial Technology Research Institute, where he developed data privacy, data de-identification, and privacy-preserving machine learning.

He is also the winner of the 2024 TAAI Best Paper Award and holds a patent related to data privacy as the first inventor.

Another co-founder of StartLux, Luo Yongxiang, is the former Managing Director of Morgan Stanley Asia.

Chen Danyan understands products and users, Guo Quanwei has long studied local models, Agents, and data privacy, and Luo Yongxiang is responsible for marketing and investment financing. This combination is highly suited to StartLux and will also help drive StartLux's long-term development.

As for what the team ultimately wants to deliver, it's not just a set of model weights.

In StartLux's vision, local intelligent solutions should be deployable with one click, similar to installing Office. Users do not need to understand quantization, GPU memory configuration, and inference frameworks, nor do they need to optimize them repeatedly themselves.

The team currently plans to launch its first-generation local smart solution for enterprise and individual users within the year.

In this light, the emergence of StartLux is by no means just "another player" entering the scene.

It also represents that, following the emergence of cloud-based model companies like DeepSeek and Kimi, domestic local models are also starting to fill in the gaps.

From the rising stars of the smart era to the first-generation programmers of the internet era, the basic paradigm of models is shifting, but Chinese companies have always been at the forefront, carrying the torch forward.

The official website link is: https://startlux.com/