What AlphaGo’s Move 37 Reveals About the Future of AI Reasoning
One afternoon in Seoul in March 2016, I watched the program I helped create place a stone on the fifth line of a Go board—a move that appeared to give its human opponent an advantage. The 37th move in the second game of the best-of-five match looked so ridiculous that some commentators thought it was a programming glitch. It was not.
AlphaGo ultimately won the match, defeating one of the greatest professional Go players of all time, Lee Sedol, by four games to one. “AlphaGo is based on probability calculations and I thought it was just a machine,” Lee said afterward. “But I changed my mind after seeing this. Yes, AlphaGo is creative.”
Why Go Required More Than Brute-Force Computing
When Deep Blue defeated Garry Kasparov, then the reigning world chess champion, in 1997, it used rules hard-coded by humans. The system could read six to eight moves ahead for each player and evaluate 200 million chess positions per second.
Go is far more complex in a different way. The value of a stone depends on how it develops across distant groups and territories over dozens of moves. It would take billions of years for a supercomputer to calculate even a fraction of the possible outcomes. To win, AlphaGo needed to recognize who had the upper hand at a glance—and invent moves that human players had never considered.
Move 37 Was Reasoning, Not Machine Intuition
This is why many accounts of the AlphaGo and Lee Sedol match describe Move 37 as a flash of pure mechanical intuition. That interpretation is misleading. The move was possible because AlphaGo combined an intuitive prediction with a deliberate search through possible futures.
Those capabilities are still missing from many AI systems today. If future AI is to produce reliable results and genuinely novel insights in areas such as science and medicine, it will need a more robust form of reasoning.
How AlphaGo Combined Intuition and Search
AlphaGo consists of two main systems. The first is a policy network trained to predict the moves a strong human player would make. Its “intuitive” assessment suggested that Move 37 was unremarkable: the probability of a skilled human choosing it was approximately one in 10,000.
AlphaGo selected the move because its search mechanism evaluated the move’s future impact beyond its immediate plausibility. It explicitly built and searched a game tree containing thousands of branches, with each branch representing a different possible future.
This resembles the theory of human thinking that Daniel Kahneman popularized. System 1 is fast, intuitive, and effortless. System 2 is slower, more deliberate, and more analytical.
AlphaGo combined both modes. Its neural network provided a hunch—a position looked promising or likely to win—while its search process tested that hunch against subsequent moves and countermeasures. Neither component would have worked as well alone. Intuition would probably never have selected Move 37, while brute-force search would have struggled to evaluate the enormous number of possible moves without guidance.
Why Today’s AI Models Still Fall Short
This architecture differs significantly from the way many AI models operate today. A large language model repeatedly selects the next token. In effect, System 1 is doing the work: producing fast, associative, and remarkably effective pattern completion across nearly every subject people write about.
Shortly after ChatGPT’s debut, the field recognized that language fluency alone was not enough for true utility. The obvious response was to create more reflective models. Instead of answering immediately, a model can break a problem into steps and generate intermediate results that influence later reasoning. This process is commonly known as chain-of-thought reasoning.
Chain-of-thought techniques have proved effective in practice, particularly in mathematics and coding. However, unlike AlphaGo’s search, they do not necessarily introduce a separate inference mechanism. The intermediate inferences are still generated through the same next-token prediction process; the model simply continues that process for longer before committing to an answer.
Three Missing Features of Reliable Machine Inference
Three shortcomings prevent chatbot behavior from qualifying as inference of the kind scientists would recognize.
- No persistent cognitive state: These models typically do not maintain an explicit, persistent, and testable record of the hypotheses they are considering, their confidence in competing explanations, the evidence they have reviewed, or the questions that remain unanswered. Such information would need to be systematically revised as new evidence arrived.
- No clear separation between knowledge and reasoning: Knowledge and reasoning are tightly intertwined in neural-network weights. There is no distinct, clearly articulated set of beliefs that the system can inspect and manipulate.
- Unreliable explanations: Although chatbot chains of thought may appear to show deliberation, research suggests that models can generate explanations that do not accurately reflect how they arrived at an answer. They may make up something or provide a different explanation after reaching the answer.
This matters because, in high-stakes applications such as medicine, engineering, and scientific research, the conclusion is not the only thing that matters. We also need to know how the system reached it. If a medical diagnosis or treatment decision is wrong, for example, we need to identify what failed. Was the reasoning flawed? Was the evidence invalid? Did the system make an incorrect assumption?
Building AI Systems That Can Track Their Reasoning
This is why I recently left my job at Google DeepMind. We believe machine inference needs a new approach—one that draws on the architecture behind AlphaGo.
AlphaGo maintains a record of what it knows about a particular position in the form of a game tree. This data structure contains the variations the system has considered, the possible futures, and the moves and positions evaluated by the neural network. As inference progresses, AlphaGo updates the tree and synthesizes the information within it to decide which move to make.
General reasoning could work in a similar way. A reasoning system would maintain an epistemic state that records what it considers resolved, what it doubts, what it has ruled out, and which questions remain unanswered.
Reasoning can therefore be understood as a sequence of actions that changes cognitive states in order to advance knowledge and reduce uncertainty. That process includes predicting possible outcomes, breaking problems into parts, and—most importantly—deciding which questions to ask, calculations to perform, or experiments to conduct next.
From Game-Tree Search to Scientific Discovery
Open-world deduction is obviously more difficult than playing a board game such as Go or chess. In the real world, the current state is only partially known, the available actions are numerous and constantly changing, and the outcomes of actions may be probabilistic or unknown.
Nevertheless, recent advances in large language models and other neural systems make it possible to address some of these inference challenges. An LLM can suggest ways to approach a problem based on the available knowledge and resources. It can also interact with tools through an API or code to help assess whether claims are supported by available evidence.
Most importantly, independent components of the system must evaluate each step according to how effectively it reduces uncertainty. Beliefs should be updated only when the change is supported by evidence. Under these conditions, a model could accumulate qualified knowledge and improve its inference policies by learning from previous reasoning experiences.
Such systems could be understood as extensions of the scientific method: machines designed to produce knowledge that can withstand scrutiny.
Why Scaling Intuition Is Not Enough
I do not think we will achieve reliable machine intelligence simply by making System 1 bigger. Scale can sharpen intuition, but it does not make intuition more thoughtful.
Move 37 matters because AlphaGo held its position, considered possible futures, and selected a move that its artificial instincts would probably have rejected. Society needs this kind of creative reasoning in fields such as drug discovery, materials science, climate research, and medical diagnostics—domains where the board does not resemble a Go board and no one has given us the rules.
These insights can come only from systems that genuinely reason: systems whose conclusions emerge from auditable evidence, inferences, and revisions to their beliefs, rather than from compelling stories generated after the fact.
Tor Grepel is Head of Machine Learning at University College London. He is a core member of DeepMind’s AlphaGo team, working to ensure that AI benefits human flourishing.
Source: www.technologyreview.com


