If you believe this, artificial intelligence models powered by thousands of advanced computer chips are intelligent. Let’s simplify concepts for toddlers.
Although infants can’t code, solve complex equations, or engage in philosophical debates, they learn about the world with incredible efficiency—unlike today’s AI models that rely on vast amounts of training data and consume substantial energy. Babies can recognize new objects after seeing them once or twice, using brief observations and physical interactions to expand their understanding rapidly.
The unique brain structures of babies hold crucial insights for enhancing AI. Developing a more baby-like AI could minimize costs and energy requirements for advanced models while enabling robots with AI to learn about their environments in a more organic manner.
In a groundbreaking initiative, researchers from Meta, Stanford University, the University of Tokyo, and France’s Ecole Normale Supérieure have developed a test that assesses infants’ learning capabilities, inspiring AI researchers to create algorithms that mimic baby learning methods.
The EgoBabyVLM Challenge introduces a visual language model (VLM) that learns from both text and images, determining how well infants comprehend their surroundings. The model ingests thousands of hours of video gathered from cameras worn by babies and young children.
Interestingly, when exposed to this authentic, chaotic footage, state-of-the-art models struggle significantly. This indicates a fundamental difference in the design of babies’ brains, allowing them to learn rapidly from minimal information.
Unlike curated datasets, infants gather knowledge from a diverse array of experiences. Parents might discuss objects that aren’t currently present, use gestures or glances to indicate things, or reference past and future events rather than focusing solely on the immediate moment. “Babies learn not only through language but also via rich multimodal experiences,” states Michael Frank, a Stanford cognitive scientist specializing in language acquisition and co-developer of EgoBabyVLM.
This research highlights that when it comes to AI, “there’s evidently more to it” than just language, according to Frank.
Language Learning
EgoBabyVLM exemplifies how researchers leverage AI to delve into human intelligence. The recently launched BabyLM challenge necessitates an AI model to learn a language’s syntax using an amount of data typical for a 10-year-old (tens of millions of words, compared to trillions for AI). Remarkably, Transformer-based AI models, which analyze language by evaluating word relationships across sentences, perform quite well. This aligns with Noam Chomsky’s theories on how syntax is ingrained in the human brain.
Ryan Cotterell, a linguist at ETH Zurich and the original creator of BabyLM, notes that the scenario differs for comprehending the physical world. “We won’t have vast collections of human interactions nor an ‘internet’ of those interactions,” he asserts.
Joshua Tenenbaum, a cognitive scientist at MIT, emphasizes that BabyLM’s results reveal the absence of “common sense” regarding physical contexts, social dynamics, and theory of mind.
“Transformers excel at pattern recognition within data,” Tenenbaum explains. “However, pure pattern-learning systems alone appear insufficient to process the type of data that infants receive and learn from.”
Source: www.wired.com


