Published October 6, 2026. Scores are as reported in the Reddit thread; check the Kaggle leaderboard for current numbers.
What happened on ARC-AGI-3?
Top scores in the ARC Prize 2026 Kaggle competition reportedly rose from about 7% to 56% within 30 days, using small models that run offline. ARC-AGI-3 was built to show tasks that are easy for people and hard for AI, so a fast jump by small local models set off a debate: real progress, or competitors learning the benchmark?
- The benchmark: interactive puzzle games where the agent must figure out the rules by playing.
- The constraint: Kaggle submissions run in a sandbox with no internet, so no GPT, Claude or Gemini API calls.
- The jump: about 7% to 56% in a month, per the thread.
- Entry deadline: October 26, 2026 (ARC Prize docs).
The Reddit thread: r/MachineLearning: “Top ARC-AGI-3 scores on Kaggle just went from 7% to…”.
What is ARC-AGI-3?
ARC-AGI is a series of benchmarks from the ARC Prize Foundation, started by François Chollet. Earlier versions used static grid puzzles. ARC-AGI-3 is interactive: the agent is dropped into small game-like environments with no instructions and has to explore, infer the goal and solve it efficiently (ARC-AGI-3 paper).
It measures skill acquisition, how quickly a system learns something new, rather than how much it already knows.
Why did the jump split Reddit?
| "This is real progress" | "This is benchmark fitting" |
| Small, offline models did it, so it is not just scale. | Public games let teams tune search strategies to the game style. |
| Agents that explore and test hypotheses are a general skill. | Hidden test games may not behave like the public ones. |
| Earlier ARC versions also fell faster than expected. | Every benchmark saturates once it becomes a target. |
Both sides can be right. Fast benchmark gains usually mix genuine technique improvements with fitting to the test.
How to read any AI benchmark
- What exactly is measured? Puzzle solving is not invoice processing.
- Is the test set hidden? Public tests get overfit.
- What resources were allowed? Compute limits, internet access, number of attempts.
- Is it independently verified? Self-reported scores deserve less weight.
- How old is it? Benchmarks saturate in months now.
The only benchmark that matters for your business
A leaderboard score cannot tell you whether an AI will answer your customers correctly or qualify your leads. The useful test is small and specific: take 20-50 real examples of the task, run the AI on them, and have the person who does the job today grade the output.
Where Dooza fits
This is how Dooza pilots work: we run the AI on your real work, measure it against how your team does the job, and you decide with evidence instead of leaderboards.
Dooza is an AI-native company that builds AI products and services for small businesses, from the Dooza Workforce app to the Dooza Agents platform. A Dooza engineer scopes your pilot on a free 30-minute call, and every product starts with a refundable pilot: 100% refund within 14 days. Book a free pilot call or see pricing.
Ready to Start Your Pilot?
Automate your business with AI employees that work 24/7. Start with a refundable pilot: 100% refund within 14 days.