Article: Fei-Fei Li and Yunzhu Li told a recent a16z podcast that World Labs is rolling out a “real-to-sim-to-real” pipeline that lets robots learn spatial tasks in high-fidelity digital twins before they ever touch a physical floor. The approach cuts the cost and risk of gathering robot training data, a hurdle that has kept many automation projects stuck in the lab.

Why robots need more than language

Large language models can scrape the entire internet for text, but a robot can’t simply watch videos and copy what it sees. To move through a warehouse or clean a hotel room, a robot must understand three-dimensional geometry, surface properties, and the physics of pushing, lifting and stacking. Those are the ingredients of “spatial intelligence,” and they are hard to capture at scale in the real world.

The data bottleneck

Collecting real-world robot data is slow, expensive and sometimes dangerous. A single mishap can damage hardware, halt a production line, or even injure a human worker. Because of that, most robotics teams rely on small, hand-crafted datasets that don’t cover the diversity of everyday environments.

World Labs sidesteps the bottleneck by building aligned digital worlds. Their SceniX system maps a physical space into a simulation that preserves geometry (the shape of walls and objects), appearance (textures, colors, lighting) and physics (how objects move when nudged). The robot practices in that virtual copy, then the same policies transfer back to the real setting.

Video models aren’t enough

Some researchers argue that high-resolution video generators could stand in for physical simulation. Those models can produce plausible frames of a robot’s viewpoint, but they lack consistent physics. If a simulated video shows a cup disappearing after a push, a robot trained on that footage will learn the wrong cause-and-effect relationship. World Labs’ pipeline insists on worlds that stay coherent across time, space and interaction, ensuring that the robot’s “muscle memory” matches real-world physics.

Targeting semi-structured spaces

The team isn’t aiming for humanoid assistants that stroll through a living room tomorrow. Instead, they focus on semi-structured environments where the layout is predictable enough for a digital twin to be accurate, yet varied enough to demand genuine spatial reasoning. Examples include:

  • Warehouses where items are stored on shelves and pallets
  • Restaurants with tables, chairs and moving staff
  • Hotels with corridors, rooms and service carts
  • Industrial sites with machinery, conveyors and safety barriers

These are the places where businesses already ask for automation. A recent survey cited in the podcast found that one-third of requested robot tasks involve cleaning—repetitive, unpleasant work that humans are eager to hand over to machines.

What success looks

World Labs’ roadmap hinges on proving the pipeline in specific industries over the next two years. If a robot can repeatedly move a pallet in a simulated warehouse and then do the same in a real facility without retraining, the model has demonstrated “reliable simulation leads to reliable robots.” That reliability could unlock larger contracts and justify the upfront investment in digital twin creation.

What to watch next

  • Industry pilots – The first handful of deployments in warehouses or hotels will reveal whether the pipeline can handle the messy edge cases that only appear on the shop floor.

The core idea is simple: give robots a perfect rehearsal space before they step onto the stage. If World Labs can turn that rehearsal into consistent performance, the hardest problem in robotics—teaching machines to understand and act in the physical world—may finally have a practical solution.