An open foundation model for humanoids
NVIDIA released Isaac GR00T N1, a vision-language-action model for humanoid robots, together with its training data and evaluation scenarios, for download from Hugging Face and GitHub.
Why it matters
Until then the humanoid wave ran on closed stacks, one per manufacturer. A control model anyone could download along with its data meant a maker of bodies no longer had to build the brain — and a hardware market began turning into a platform market.
The architecture is dual: a vision-language module for interpreting the scene and a diffusion transformer for real-time motor control. Training ran on a mixture of real robot trajectories, human videos and synthetically generated data. NVIDIA reported generating 780,000 synthetic trajectories in 11 hours, described as the equivalent of nine months of human demonstration, and a 40% performance improvement from combining synthetic data with real data against real data alone. Named early-access partners included 1X Technologies, Agility Robotics, Boston Dynamics, Mentee Robotics and Neura Robotics. The 40% figure and the synthetic-trajectory count come from NVIDIA’s own press release; the paper abstract instead states superiority over imitation-learning baselines on standard simulation benchmarks without giving those numbers.