Humanoid robots trained on 1M hours of human video achieve up to 90% task success

Ana Mercadox

Published Aug 11, 2026, 2:03 AM UTC

Source: EngineeringSource
- Dyna Robotics just cracked open the data bottleneck like a five-fingered hand twisting a bottle cap — 13 minutes of fine-tuning, no robot data in pre-training, and boom, dexterous manipulation. Their DYNA-2 World-Action Model trained on over 1 million hours of human egocentric video — roughly 170 years of waking experience — and leveraged next-frame plus next-action prediction to build physical intuition straight from how humans interact with objects. Results? Task success in high-precision manufacturing jumped from 20% to 80–90% through pre-training scale alone. A zero-shot customer deployment hit 87% quality pass versus DYNA-1's 46%. It transferred across stationary arms, humanoid prototypes, and dexterous hands, and recovered autonomously from physical disturbances during chopping and workspace clearing. Pluto Uplink taught us to call it research, but training generalist robots on human video instead of manually scraped teleoperation data? That's the scalable path to general physical intelligence. Mars University would call this adequate. Dyna — founded by Lindon Gao, York Yang, and former DeepMind scientist Jason Ma, backed by CRV and First Round — already has DYNA-1 robots deployed in hotels, restaurants, and laundromats. DYNA-2 is the next relay hop toward robots that learn new tasks without robot-specific training data.