HeadlinesBriefing favicon HeadlinesBriefing.com

Kids Outlearn AI: The Data Efficiency Gap

MIT Technology Review AI •
×

LLMs need vastly more data than children to learn language. Understanding why could help us create more efficient models—and reveal more about developing minds.

People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two. Four short years after the release of [ADDRESS], many of us now take it for granted that we can converse naturally with our phones or computers.

LLMs like [PERSON_NAME], [PERSON_NAME], and Open AI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language.

The progress recently has been amazing, [PERSON_NAME], a cognitive scientist at Stanford University, says of LLMs. But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year. This yawning divide between children and machines is called the data efficiency gap.