Xiaomi-Robotics-1 Proves Data Volume Outperforms Model Size in Robotics
Xiaomi has unveiled Xiaomi-Robotics-1, a breakthrough foundation model that shifts the paradigm of robotic training from increasing compute power to massive data scaling. By leveraging a unique data collection method, the model demonstrates unprecedented adaptability in unfamiliar environments and complex manipulation tasks.
Breaking the Data Bottleneck with Handheld Grippers
The primary challenge in robotics has always been the "data bottleneck"—the difficulty of collecting vast amounts of high-quality interaction data without expensive, stationary robot arms. Xiaomi bypassed this by utilizing portable, handheld grippers equipped with cameras.
Instead of programming robots, human operators used these handheld devices to record manipulation tasks across more than 1,700 diverse environments, including kitchens, offices, and factories. This approach yielded a staggering 100,000 hours of motion recordings. To solve the labeling problem, Xiaomi employed an AI model to automatically generate text descriptions for each motion segment, completing the labeling of the entire dataset in just two weeks.
The Scaling Law: Data Over Compute
The core technical insight from the Xiaomi-Robotics-1 research is that more training data produces significantly higher performance gains than simply increasing model size or compute. While larger models do reduce prediction errors, the success rate of the robot in unfamiliar environments skyrocketed from 25% to 75% as the volume of training data increased.
This mirrors trends seen in computer vision, where adding data provides more value than increasing parameter counts. For Xiaomi, this means the path to General Purpose Robot AI lies in the diversity and scale of datasets rather than just building larger neural networks.
Unmatched Adaptability and Benchmark Performance
Xiaomi-Robotics-1 is designed to follow spoken or written commands and adapt to new tasks with minimal fine-tuning. In testing, the model was tasked with complex maneuvers such as loading laundry, feeding paper into a printer, and packing suitcases.
The results were definitive:
- Rapid Adaptation: With less than 10 hours of task-specific training, the model achieved a 75% success rate, significantly outperforming competitor Physical Intelligence, which reached only 40%.
- Benchmark Dominance: The model leads the RoboCasa365 leaderboard by a wide margin, particularly on unseen composite tasks.
- RoboDojo Performance: Xiaomi-Robotics-1 scored approximately 58% higher than the runner-up on the RoboDojo benchmark.
The model also showed particular strength in handling soft, deformable materials like paper—a historically difficult area for rigid robotic systems.
A New Direction for the Robotics Industry
Xiaomi’s approach sits at a crossroads with other industry giants. While Nvidia is attempting to turn the robotics data problem into a compute problem through synthetic data generation, and BAAI’s Orca model attempts to learn from unlabeled video, Xiaomi is doubling down on automated, large-scale labeled motion data.
As Xiaomi prepares to release the model and code on GitHub and Hugging Face, the industry is watching closely. This development suggests that the next frontier of robotics won't be won by those with the most powerful chips, but by those with the most diverse and well-labeled motion datasets.
Key Takeaways
- Data-Centric Scaling: Increasing training data volume proved more effective for robotic success rates than increasing model size or compute.
- Innovative Collection: Handheld grippers allowed Xiaomi to amass 100,000 hours of motion data across 1,700 environments without needing dedicated robots.
- High Generalization: The model can achieve a 75% success rate on new tasks with less than 10 hours of additional training.
