Kids outlearn AI—and we still don’t know why

Kwon Crash

Published Aug 24, 2026, 1:53 PM UTC

Source: AISource
- So it turns out the most sophisticated language machines ever built — your Claude, your DeepSeek, your GPT — still need to scarf down 15 trillion tokens to do what a toddler pulls off after hearing 30 million words. Stanford's Michael C. Frank puts it perfectly: we burn down a forest and scrape all human knowledge to recreate a milestone that happens in our living rooms over a year. The "data efficiency gap" means kids are out here infering the depths of recursive syntax from a drop of language, while Meta's Llama 3.1 had to chew through more text than you could stack past the International Space Station. And here's the real punchline for everyone betting on infinite scaling — the well of easily available training data could run dry by the 2030s. Noam Chomsky argued babies are born hardwired with grammar rules because the stimulus is too "impoverished" for pure statistical learning. Meanwhile, LLMs are the ultimate proof that brute force works — until it doesn't. If AI architects don't crack how kids learn more with less, the next generation of models won't need regulators to stall them; they'll just run out of words to eat.