What Is a Humanoid Foundation Model? An Introduction to GR00T N1 - Annika & Aastha
Jul 28, 2025 · 17:47
NVIDIA's Annika Brundyn and Aastha Jhunjhunwala introduce GR00T N1, an open-source Vision-Language-Action foundation model for humanoid robots. They argue that physical AI is key to addressing labor shortages in industries like healthcare and manufacturing, and that humanoid forms are needed because the world is built for humans. The model uses a data pyramid strategy combining limited real-world teleoperation data, synthetic simulation data, and internet video, with DreamGen for data multiplication. Its dual-system architecture, inspired by Kahneman's 'Thinking, Fast and Slow', has a System 2 planner for high-level reasoning and a System 1 for fast motor control at 120 Hz. The model is trained end-to-end via imitation learning and reinforcement learning, and features a generalist action decoder that enables cross-embodiment fine-tuning.