“It’s an absolute miracle,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator—but you don’t get a child.”
How babies learn language remains one of the biggest mysteries in developmental psychology and linguistics. Researchers understand much about what children learn and how their language skills develop, but the fundamental question—why humans can acquire language so quickly—is still unanswered.
Human language is built on syntax: a complex system of rules for combining words into sentences. These rules include recursive and nested structures, allowing people to express virtually unlimited ideas with a limited number of words and sounds. For babies, this presents an enormous challenge. They are immersed in only a small amount of speech, yet they somehow infer the deeper structure of language from it. They estimate the depth of the ocean from its droplets.
One influential explanation was proposed by MIT linguist Noam Chomsky in the 1950s. Chomsky argued that babies are born with an innate understanding of grammar. His theory challenged the behaviorist perspective of psychologist B.F. Skinner, who believed that language acquisition occurs primarily through environmental conditioning and reinforcement—similar to the way a dog learns to sit or shake hands in exchange for a treat.
Chomsky countered with the concept of the “poverty of the stimulus.” The theory suggests that children’s exposure to speech is too limited and imperfect to explain how they acquire such a sophisticated understanding of grammar through experience alone. “His signature argument was essentially that language cannot be learned purely on a statistical basis,” said Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Chomsky instead proposed that language depends on logical rules and that children possess innate knowledge that helps them infer grammar from incomplete speech.
“It’s an absolute miracle… If you train GPT-2 on 30 million words, you get a nonsense generator, but you don’t get a child.”
Michael C. Frank, cognitive scientist, Stanford University
Chomsky’s theory dominated American linguistics for decades through the framework known as generative grammar. His ideas also influenced the development of artificial intelligence and computer science during the 1950s and 1960s. At the beginning of the AI boom, linguistics and natural language processing became closely connected, supported by extensive military funding. The Pentagon wanted computers that could understand English and translate Russian.
Although early neural networks showed that machines could learn and reproduce statistical patterns, many American AI researchers favored rule-based systems influenced by Chomsky’s theories. Their goal was to teach computers language by explicitly programming grammatical rules. Instead of exposing machines to vast amounts of natural speech, researchers focused on formal grammar lessons. This approach became part of symbolic AI, a movement that remained influential for decades but largely failed to produce systems capable of understanding human language at scale. Interest in natural language processing declined during the AI winter that began in the 1970s.
Neural networks eventually returned to prominence. However, it was not until the 2010s—when computing hardware became more powerful and affordable and internet data became widely available—that their capabilities attracted widespread attention. By 2018 and 2019, language models such as BERT and GPT-2 demonstrated the effectiveness of training on enormous datasets. Built on the Transformer architecture and trained on billions of tokens, these models showed that statistical learning could produce surprisingly sophisticated language abilities. In 2022, OpenAI’s ChatGPT brought large language models into the mainstream.
Large language models are not human brains. They are powerful statistical learning systems: sophisticated pattern-recognition machines without the biological development, sensory experiences, or evolutionary features of the human cortex. In many ways, they appeared to accomplish what generative linguists had once argued machines could not do—learn the structure of language. These models can now write convincing sonnets, generate fluent prose, and perform well on grammar tests.
“No matter how skeptical you are about AI, what impresses everyone is that these systems learn syntax,” said Alison Gopnik, a developmental psychologist at the University of California, Berkeley. “I didn’t think that would be true. And I don’t think most people thought you could understand grammar simply by analyzing the statistical patterns in a large sample of language.”
Source: www.technologyreview.com


