Technology Trends

ARC-AGI-3 Scores Jumped From 7% to 56% on Kaggle. What That Does and Does Not Mean

Small offline models reportedly went from about 7% to 56% on ARC-AGI-3 in a month. Here is what the benchmark tests, why the jump caused an argument on Reddit, and how to judge any AI benchmark for your own work.

6 min read
October 6, 2026
Watercolor illustration of a small robot solving a colored grid puzzle next to a rising line chart

Published October 6, 2026. Scores are as reported in the Reddit thread; check the Kaggle leaderboard for current numbers.

What happened on ARC-AGI-3?

Top scores in the ARC Prize 2026 Kaggle competition reportedly rose from about 7% to 56% within 30 days, using small models that run offline. ARC-AGI-3 was built to show tasks that are easy for people and hard for AI, so a fast jump by small local models set off a debate: real progress, or competitors learning the benchmark?

  • The benchmark: interactive puzzle games where the agent must figure out the rules by playing.
  • The constraint: Kaggle submissions run in a sandbox with no internet, so no GPT, Claude or Gemini API calls.
  • The jump: about 7% to 56% in a month, per the thread.
  • Entry deadline: October 26, 2026 (ARC Prize docs).

The Reddit thread: r/MachineLearning: “Top ARC-AGI-3 scores on Kaggle just went from 7% to…”.

What is ARC-AGI-3?

ARC-AGI is a series of benchmarks from the ARC Prize Foundation, started by François Chollet. Earlier versions used static grid puzzles. ARC-AGI-3 is interactive: the agent is dropped into small game-like environments with no instructions and has to explore, infer the goal and solve it efficiently (ARC-AGI-3 paper).

It measures skill acquisition, how quickly a system learns something new, rather than how much it already knows.

Why did the jump split Reddit?

"This is real progress""This is benchmark fitting"
Small, offline models did it, so it is not just scale.Public games let teams tune search strategies to the game style.
Agents that explore and test hypotheses are a general skill.Hidden test games may not behave like the public ones.
Earlier ARC versions also fell faster than expected.Every benchmark saturates once it becomes a target.

Both sides can be right. Fast benchmark gains usually mix genuine technique improvements with fitting to the test.

How to read any AI benchmark

  1. What exactly is measured? Puzzle solving is not invoice processing.
  2. Is the test set hidden? Public tests get overfit.
  3. What resources were allowed? Compute limits, internet access, number of attempts.
  4. Is it independently verified? Self-reported scores deserve less weight.
  5. How old is it? Benchmarks saturate in months now.

The only benchmark that matters for your business

A leaderboard score cannot tell you whether an AI will answer your customers correctly or qualify your leads. The useful test is small and specific: take 20-50 real examples of the task, run the AI on them, and have the person who does the job today grade the output.

Where Dooza fits

This is how Dooza pilots work: we run the AI on your real work, measure it against how your team does the job, and you decide with evidence instead of leaderboards.

Dooza is an AI-native company that builds AI products and services for small businesses, from the Dooza Workforce app to the Dooza Agents platform. A Dooza engineer scopes your pilot on a free 30-minute call, and every product starts with a refundable pilot: 100% refund within 14 days. Book a free pilot call or see pricing.

Frequently Asked Questions

What is ARC-AGI-3?

An interactive AI benchmark from the ARC Prize Foundation in which agents must learn the rules of unfamiliar game-like environments by exploring them. It measures how efficiently a system acquires new skills.

What is the top ARC-AGI-3 score on Kaggle?

A viral r/MachineLearning thread in early October 2026 reported top Kaggle scores rising from about 7% to 56% within 30 days. Check the Kaggle leaderboard for current figures.

Can ARC Prize Kaggle entries use GPT or Claude?

No. Kaggle submissions run in a sandbox without internet access, so they cannot call hosted model APIs. Teams use models that run offline.

Do AI benchmark scores predict business results?

Not reliably. Test AI on 20-50 real examples of your own task and have the person who does that job grade the results.

Ready to Start Your Pilot?

Automate your business with AI employees that work 24/7. Start with a refundable pilot: 100% refund within 14 days.

Related Articles

Is Your AI Chat Private? What the Claude Diary Arrest Means for You and Your Business
Technology Trends

Is Your AI Chat Private? What the Claude Diary Arrest Means for You and Your Business

A Florida woman treated Claude like a private diary. A threat she wrote was flagged, reviewed by people at Anthropic and sent to police. Here is what AI chats are and are not, and the policy every business using AI should write this week.

7 min read
Read
PewDiePie, Ajax and the OpenAI Bans: What AI Distillation Is and Why It Gets You Banned
Technology Trends

PewDiePie, Ajax and the OpenAI Bans: What AI Distillation Is and Why It Gets You Banned

PewDiePie says OpenAI banned him twice while he trained Ajax, a 9B local model, on outputs from OpenAI models. Here is what distillation means, why AI labs ban it, and the platform risk lesson for businesses.

7 min read
Read

Ready to scale your business?

Start with a refundable pilot — 100% refund within 14 days. A Dooza engineer scopes it with you on a free 30-minute call. Pricing depends on the product; see pricing.

Refundable pilot · 100% refund within 14 days · No contracts