A baby learns to talk from a trickle of overheard speech. Feed a language model the same amount of text and you get gibberish. That gap sits at the center of a new piece in MIT Tech Review, which digs into a question researchers still can’t answer: why can kids learn language at all, when the machines built to do the same thing need thousands of times more data?
The number that makes this concrete comes from Stanford cognitive scientist Michael C. Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid,” he told MIT Tech Review. Roughly 30 million words is what a child hears in the first few years of life. It’s enough for a toddler to master grammar. It’s nowhere near enough for a transformer.
The old fight AI just walked into
This debate is older than the technology. Back in the 1950s, MIT linguist Noam Chomsky argued that babies are born with hardwired knowledge of grammar. He was pushing against B.F. Skinner, who thought language got learned the way a dog learns to sit, through reward and repetition. Chomsky’s counter was the “poverty of the stimulus”: kids simply don’t hear enough language, and what they hear is too messy, to reverse-engineer its rules from scratch. Something had to be built in.
As UC Irvine linguist Richard Futrell put it to MIT Tech Review, Chomsky’s core claim was “that language cannot be learned on the basis purely of statistics.”
That idea shaped decades of AI. Flush with Cold War funding to translate Russian and parse English, US researchers tried to teach computers language by hand-coding the rules. Grammar class, not immersion. This symbolic approach dominated, and it largely failed. Natural-language processing went cold in the “AI winter” of the 1970s.
Then the statistics won
Neural networks crept back as hardware got cheap and the internet got huge. By 2018 and 2019, BERT and GPT-2, built on the transformer architecture and trained on billions of tokens, showed insiders that raw scale could crack language. ChatGPT made it obvious to everyone in 2022.
Here’s what stands out. Large language models are not brains. They’re naive pattern machines, none of the evolved biology baked into the human cortex. They are, in Chomsky’s own framing, exactly the kind of thing that should not be able to learn language. And they did it anyway.
“No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax,” UC Berkeley developmental psychologist Alison Gopnik told MIT Tech Review. “I didn’t think that was going to turn out to be true.”
Why this matters for anyone building with AI
The headline isn’t “AI beat the babies.” It’s the opposite, and the practical read is worth holding onto:
- Data efficiency is the real frontier. A child hits fluent grammar on 30 million words. Frontier models need billions. If you care about training cost, the human brain is proof that a far leaner path exists. We just don’t know the recipe.
- Scale solved a problem theory said was unsolvable. For decades the expert consensus was that statistics alone couldn’t produce grammar. That consensus was wrong. Worth remembering the next time someone tells you a capability is off the table.
- “It works” and “we understand it” are different things. Models learn syntax, and nobody fully knows how babies do, or how the models do either. Two black boxes, both delivering, neither explained.
The honest limitation
The researchers aren’t claiming victory for either side. Models passing grammar tests doesn’t prove Chomsky was wrong about babies, and it doesn’t prove Skinner was right. Kids and machines could be reaching the same destination by completely different roads. As the MIT Tech Review piece makes plain, the mechanism inside a toddler’s head is still, in Frank’s word, “miraculous,” and still unexplained.
What comes next is the interesting part: researchers are now using these models as testbeds for old theories about the human mind. The full story is worth reading at MIT Tech Review.