Video Game Data Could Help Train AI to Navigate the Real World
Round and round By moving a thumbstick, squeezing a trigger, and pressing a few buttons, even unskilled players can navigate complex 3D video game environments. A British startup is betting that these simple sequences contain a valuable source of data for training new artificial intelligence models.
Why AI Needs More Than Language
Some researchers in the AI industry believe large language models may eventually be limited by their inability to navigate the physical world. Because LLMs are trained primarily on words, they may be poorly equipped to drive a car, maneuver a robotic arm, or perform other tasks requiring delicacy and precision. To address this limitation, researchers including Feifei Li and Yang Lukun are focusing on a different type of artificial intelligence: a world model.
To understand real-world physics, world models must be trained on both visual and action data. Before an AI system can dexterously operate a robotic arm, for example, it may need video footage from a factory floor combined with information about how firmly to grip an object and how much torque to apply. Unlike the vast collections of text used to train LLMs, however, there is no comparable global archive of data for training world models.
“World models need cause and effect,” says Xiatian Zhu, an associate professor of AI at the University of Surrey. “There’s very little data of this kind on the internet.”
Worldmodeldata Wants to Turn Games Into AI Training Data
Worldmodeldata, a British startup advised by LeCun, aims to address this shortage by packaging controller inputs and other information collected by video game studios into datasets for training world models. Some companies, including General Intuition and Niantic, are already collecting video game data from their platforms to build AI models. Worldmodeldata is positioning itself as a broker that curates and organizes this information, allowing AI labs to avoid negotiating individual agreements with dozens of game studios.
“There are millions of great games, and they are becoming more and more like the real world,” Worldmodeldata CEO Leah Lucas told WIRED. “Why not take the vast, rich, diverse experiences from video games and teach them to AI?”
Why Video Games Could Offer Valuable World-Model Data
The idea has not yet been fully tested, but researchers have studied it based on the assumption that, much like LLMs, world-model performance improves as the size of the training dataset increases. That makes the lack of high-quality training data one of the biggest potential bottlenecks for progress.
Some AI labs are trying to generate their own data by attaching sensors to people and robots in controlled test environments. However, this approach produces relatively small amounts of information and may fail to capture the unusual situations a model could encounter in the real world.
“You can pay people to demonstrate pick-and-place tasks, but just repeating them doesn’t capture the disorder in the world that we’re asking machines to perform,” says Nicole Frankel, a partner at Khosla Ventures, a venture capital firm that has invested in General Intuition.
Worldmodeldata’s theory is that data from video game environments—visual representations of 3D spaces combined with the actions players take—is available at the necessary scale and diverse enough to capture important edge cases.
“It’s the special cases that you can actually get right,” Frenkel says. “The cost of error in cars, airplanes, drones, factory forklifts, and autonomous quadrupedal vehicles is very high.”
Nearly 1 Million Hours of Licensed Game Data
Worldmodeldata says it has licensed nearly 1 million hours of data from several popular video game studios, although Lucas declined to identify them. The startup also hopes to create a system that allows individual players to receive compensation for contributing data in the future.
Source: www.wired.com


