Why “Recursive Intelligence” and AI Reasoning Claims Need Better Evaluation
Of course it is!
Is Recursive Intelligence in AI Real?
The New York Times ran a major article about how scientists and researchers are now focusing on recursive intelligence—AI that learns from other AI systems. Is that a real development? It sounds like you’re saying it isn’t.
AI researchers have a problem: the field often uses ambitious names for ambitious ideas. “Machine learning” is an ambitious name.
But machine learning is real.
The name can still be more ambitious than the technology itself. The machine is not always there.
When AI Terminology Overstates What Models Do
Just because AI has curriculum learning does not necessarily mean it represents a distinct, fully established field. Let me introduce you to a paper by my former manager, Sammy Bengio. Although he is a machine learning superstar, he is not as famous as his older brother, Joshua.
I asked him, “Are you tired of living a life where every time you write a paper, you debunk your entire reasoning?”
He left Google after I was fired and is now head of machine learning research at Apple. If you look at almost every paper they have, you’ll see that changing the inference benchmark even slightly can cause the entire result to break down. That’s not logical.
Understood.
Let me give you another example. A model trained to output a specific token—a probabilistic pattern—is simply described as performing chain-of-thought inference. You did not know what the model was thinking, and you did not know that it was a chain. You only knew that it produced printed tokens. Yet it was called chain-of-thought reasoning. Now you’re saying the model has already reasoned.
The Scientific-Research Crisis in AI
One of the biggest crises we currently face is the reliability of scientific research. When I look at my own work, even though I had access to the data, the code, the training data and the evaluation data, I still do not know whether I captured a benchmark during training because these companies do not provide all of that information.
Ingesting benchmarks during training is like studying for a test. It is like entering an exam already knowing the answers to 10 questions, studying them and then taking the test. Every time I have investigated these systems, I have shown how incorrect their claims can be.
Why AI Benchmarks Can Fail
What Sammy Bengio and his team showed in their 2025 paper is that changing a benchmark slightly can cause the entire result to break down. The model may have been relying on a pattern rather than demonstrating genuine reasoning.
When I give this example, some people say, “But we’re talking about a 2026 model.”
The answer is that meaningful scientific evaluation takes time. We must not move from a press release to a lawmaker parroting that press release, or from a press release to a journalist repeating its claims.
If you want to conduct real research and evaluation, request the training data, evaluation data and methodology. That is how results can be independently reproduced.
Is AI Reasoning Still Fundamentally Different From Human Thinking?
So it seems like you still think of thinking, reasoning and recursive intelligence as purely human abilities. That’s how our neural processes work, but AI does not work that way yet. Is that what you believe?
Source: www.wired.com


