The Scaling Illusion and the Rise of a Tiny Competitor
A recursive model with 7 million parameters scored higher on the ARC-AGI-1 reasoning benchmark than several far larger LLMs, including Gemini 2.5 Pro, o3-mini-high and DeepSeek R1 (671B parameters). Its predecessor HRM, at 27 million parameters, did something similar a few months earlier. That does not mean small models beat big ones at everything. It does show that on this kind of grid puzzle, letting a small network loop over its own answer counts for more than raw size, at a tiny fraction of the compute.
This article tells the story of the emergence of Small Recursive Models that have changed the rules of the game.
While Large Language Models (LLMs) struggle with logical puzzles that we call AGI benchmarks, a new and fascinating idea has emerged against this "scaling virus." The idea is simple: What if instead of making the model's brain bigger, we allow it to think before answering and refine its response?
This idea led to HRM (Hierarchical Reasoning Model), published by Sapient Intelligence in June 2025 (arXiv 2506.21734). It has 27 million parameters and was trained on roughly 1,000 examples, with no pretraining on internet text. The HRM paper reports 40.3% on ARC-AGI-1 and 5.0% on ARC-AGI-2, ahead of o3-mini-high (34.5%) and Claude 3.7 8K (21.2%) on ARC-AGI-1. Those are public evaluation set scores. When the ARC Prize team re-ran HRM on its hidden semi-private set, it measured 32% on ARC-AGI-1 and 2% on ARC-AGI-2 (ARC Prize analysis). The same analysis found that the two-level "hierarchical" design added little over a plain transformer of the same size; most of the gain came from the outer refinement loop.
How Does the HRM Model Work? (The Secret of Slow and Fast Thinking)
The key difference is that regular language models try to predict everything in a single "Forward Pass." This means the effort they put into calculating "1+1" equals the effort for a complex problem.
But the HRM model is inspired by the structure of the human brain. This model consists of two small transformer networks that operate at different speeds:
- Low-level network (Fast): Makes quick, incremental changes on a mental "Scratchpad."
- High-level network (Slow): Determines the overall strategy and decides whether the answer is ready or needs more thinking.
Instead of instant responses, this model thinks recursively and refines its answer multiple times until it reaches the correct result.
The TRM Model: When "Small" Gets Even Smaller!
An even simpler follow-up, TRM (Tiny Recursive Model), came from Alexia Jolicoeur-Martineau at Samsung SAIL Montréal in October 2025 (arXiv 2510.04871). It drops HRM's biological assumptions, uses a single small network, and changes how training works. At 7 million parameters, about a quarter of HRM's size, it scores higher.
The TRM paper reports 44.6% on ARC-AGI-1 and 7.8% on ARC-AGI-2 (usually rounded to 45% and 8%). In the same table Gemini 2.5 Pro gets 37.0% and 4.9%, o3-mini-high 34.5% and 3.0%, DeepSeek R1 15.8% and 1.3%. It does not beat everything: Grok-4-thinking scores 66.7% and 16.0%. ARC Prize's independent re-test on the semi-private set measured 40% on ARC-AGI-1 and 6.2% on ARC-AGI-2 (ARC Prize results). GPT-4 and GPT-5 do not appear in either paper's comparison.
Why Was TRM More Successful?
Instead of mimicking the mouse brain (which was used in HRM), the TRM model focused on a functional separation:
- A space for the Thinking Scratchpad.
- A space for the Answer Placeholder.
Instead of assuming it reaches an "equilibrium" (a mistake HRM made), this model trains exactly on the loops it executes. The most interesting point is that when researchers tried to increase the model's layers, its performance decreased. In fact, the model's small size prevented it from "Overfitting" on limited data and helped it generalize better.
An Analogy for Better Understanding
Imagine asking two people to draw a complex painting.
The first person (a large language model): Is a genius who must complete the painting with one stroke of the pen without lifting their hand from the paper. No matter how genius they are, the probability of error is high.
The second person (TRM): Is an ordinary painter, but they're allowed to draw an initial sketch, look at it, erase, correct, and work on details for hours until they're satisfied with the result.
In logical and complex problems, the second painter (small recursive model) often wins because they have the opportunity to think and correct, even if they have a smaller brain (parameters).
Conclusion: The Future of AI is Recursive
This research proved that solving hard logical problems is not exclusive to giant language models. Small models with recursive thinking capability (Recursion) can solve problems that large models cannot handle with a single processing pass.
Instead of memorizing data, these models learn how to edit a Canvas over multiple steps until they reach the correct answer.
Key Takeaways from This Research:
- Bigger isn't always better - architecture and thinking method matter more than size
- The ability to review and refine can replace billions of parameters
- Inspiration from the human brain (fast and slow thinking) can lead to more efficient models
- Smaller models can generalize better and are less prone to overfitting
Perhaps the future of AI lies not in trillion-parameter models, but in intelligent models that know how to "think."
This is the same bet FanMind makes on a smaller scale: a narrower, citation-required assistant that declines to guess beats a broader one that answers everything fluently but not always correctly. See how that plays out for an internal organizational assistant.