On September 2, Dr. Sun Pengbo, a former reinforcement learning expert at ByteDance and former head of Tencent's Robotics X Intelligent Agent Center, officially joined Xingchen Intelligence, where he will focus on technical R&D and application exploration in robot reinforcement learning.
Over the past two years, robotics foundation models have largely focused on whether robots can perform a task at all. But as robots move from demos to real-world deployment and commercialization, the question is shifting to whether they can perform reliably. The need to continuously refine strategies based on actual execution results is putting reinforcement learning back at the center of industry attention.
Sun Peng has long been dedicated to reinforcement learning (RL), intelligent agents, and multi-agent reinforcement learning (MARL), with research and industrial experience spanning robot reinforcement learning, large-scale distributed RL systems, and large model reinforcement learning. His joining will further strengthen Stardust Intelligence's layout in robot reinforcement learning, and form a synergy with AI models, embodied OS, and cable-driven robot entities, driving robots to continuously learn and evolve in the real world.
After joining Xingchen Intelligent, Sun Peng continued to expand the robotics reinforcement learning team, driving the enhancement of related technical capabilities and application deployment.

From Tencent to ByteDance, delving deep into reinforcement learning for over a decade
Dr. Sun Peng graduated from Tsinghua University and subsequently conducted postdoctoral research at Cornell University and Rutgers University. Over the past decade, his research and industrial practice have consistently focused on reinforcement learning and intelligent agents, gradually expanding from robotic reinforcement learning to large-scale RL systems and post-training of large models, forming a technological accumulation that combines algorithmic research and system engineering.
During his early tenure at Tencent AI Lab and Robotics X, Sun Peng served as the head of the Intelligent Body Center, where he conducted research on deep reinforcement learning and robotic control. He utilized deep reinforcement learning and adversarial games to train wheeled robots, achieving end-to-end active target tracking. At the same time, he developed TStarBots and TStarBotX, AI intelligent bodies for StarCraft, and made significant breakthroughs in the research of complex game environments for multi-agent reinforcement learning (MARL).
After joining ByteDance, Sun Peng expanded his technical focus from algorithms to large-scale RL systems engineering. He led the development of ByteRL, the company's core reinforcement learning infrastructure, which supports efficient training for multiple in-house game AIs. His team won the IEEE CoG 2023 strategy card game AI competition, and their Hearthstone AI demonstrated competitive performance against top industry players.
With the rise of large language models, Sun Peng further expanded his accumulation of reinforcement learning to post-training of LLMs. During his time at Byte AI Lab/ByteResearch, he participated in research as a project leader and released the reinforced fine-tuning method ReFT (Reinforced Fine-Tuning) and the math reasoning intelligent body DeltaProver. At Seed, he deeply participated in model RLHF and pre-training, accumulating rich cutting-edge exploration experience in model alignment, reasoning enhancement, and Agent direction.
Sun Peng's technical path over the past decade, from robotic reinforcement learning to large-scale RL systems and post-training of large models, has always revolved around a core issue: how to enable intelligent agents to continuously learn through environmental interaction and feedback, and optimize their strategies. Now, as AI further moves from the digital world to the physical world, this capability is also finding new applications in the field of embodied intelligence.
Xingchen Intelligent Boosts Robot Post-Training Closed Loop
For Xingchen Intelligent, while advancing the "AI model + embodied OS + cable-driven entity" full-stack technology system, Sun Peng's joining is a crucial link in solidifying the robot model's "post-training" phase.
Since its inception, Xingchen Intelligent has adhered to Design for AI, believing that the robot's body, data, and AI models should not be designed in isolation. Over the past few years, this approach has gradually formed a complete closed loop covering base models and commercial models, front-end symbiotic Agents, and reinforcement learning post-training.
Xingchen Intelligent's base model series, Lumo, continues to advance robots' understanding, prediction, and generation of their environment and actions. Lumo-1 enables robots to understand the "why" behind actions, while Lumo-2, as the first implicit world action model for households, introduces Latent World Dynamics, predicting future worlds in latent space before generating actions. The front-end Agent Philia focuses on long-term memory, task management, multi-robot collaboration, and natural interaction. Reinforcement learning, as a foundational model, plays a crucial role in scaling up to real-world deployment, addressing how robots learn from continuous interaction and task feedback in real environments and optimize their behavior strategies.
Sun Peng's long-term accumulation in robot reinforcement learning, large-scale RL systems, and large model RLHF, reinforcement fine-tuning, and Agent direction will further strengthen Stardust Intelligence's technical capabilities in this area. In the future, he will participate in the construction of Stardust's embodied model, robot reinforcement learning, and post-training system, exploring how to further combine the reinforcement learning post-training methods that have been verified in the era of large models with real robot execution, data closed-loop, and long-term deployment.
Regarding his decision to join Xingchen Intelligent, Sun Peng said: "Over the past decade, I have been working on reinforcement learning and intelligent bodies, and I have witnessed the development of reinforcement learning from robot control and complex games to post-training of large models. Now, I hope to apply these accumulations to the physical world. Xingchen Intelligent already has a complete technical foundation, from robot hardware, embodied OS to AI models. What I look forward to most is working with the team to further integrate reinforcement learning into the training and deployment of real robots, enabling robots to not only 'learn a task', but also to improve their performance through repeated real-world execution and feedback."
This article is provided by Xingchen Intelligence and reproduced by QbitAI with permission, with opinions belonging to the original author.