Robotics: why now? - Quan Vuong and Jost Tobias Springberg, Physical Intelligence
Jul 26, 2025 · 18:07
Physical Intelligence's Quan Vuong and Jost Tobias Springberg describe their mission to build a model that can control any robot to do any task, arguing that software intelligence is the main bottleneck in robotics. They explain Vision Language Action models (VLAs) as adaptations of vision language models that output robot actions instead of text. To train these models, they built a data engine from scratch, collecting 10,000 hours of successful episodes via teleoperation in six months. Their latest model, PAIO-5, achieves open-world generalization by training on data from multiple homes, matching or surpassing performance on held-out scenes. They demonstrate this with a policy that performs long-horizon tasks like cleaning an unseen bedroom for up to 10 minutes autonomously. They also highlight a remote coffee-making demonstration on a robot they never touched, showing model portability across hardware.