ARC-AGI-2
The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.
ARC-AGI-2 remains the current static colored-grid ARC benchmark, but the official ARC Prize leaderboard generated on 2026-07-23 lists GPT-5.6 Sol Max at 92.5% and GPT-5.6 Sol xHigh at 90.0% on the ARC-AGI-2 semi-private axis. Because competition/private-set validation and efficiency constraints still matter, this page marks it nearing saturation rather than fully saturated.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| DeepSeek V4-Flash | 67.5% | 2026-09-12 | independent | Source ↗ |
| DeepSeek V4-Pro | 71.8% | 2026-09-12 | independent | Source ↗ |
| Claude Fable 5.1 (Hybrid Reasoning) | 70% | 2026-09-12 | independent | Source ↗ |
| Gemini 3.7 Flash | 72.5% | 2026-09-10 | independent | Source ↗ |
| Grok 4.6 | 72.9% | 2026-09-09 | independent | Source ↗ |
| GPT-5 | 73.9% | 2026-09-04 | independent | Source ↗ |
| DeepSeek V4.1 Flash | 69% | 2026-08-29 | independent | Source ↗ |
| GPT-6 Astra (max) | 71.5% | 2026-08-25 | independent | Source ↗ |
| Kimi K3 | 73.2% | 2026-08-23 | independent | Source ↗ |
| Claude Opus 5 | 74.5% | 2026-08-21 | independent | Source ↗ |
| Claude Sonnet 5 | 72.5% | 2026-08-21 | independent | Source ↗ |
| Gemini 3.5 Flash-Lite | 66.7% | 2026-08-18 | independent | Source ↗ |
| Gemini 3.6 Flash | 71.4% | 2026-08-18 | independent | Source ↗ |
| GPT-5.6 Sol | 74.9% | 2026-08-06 | independent | Source ↗ |
| GPT-5.5 Thinking | 73.7% | 2026-08-06 | independent | Source ↗ |
| MiniMax M3 | 68.6% | 2026-07-23 | independent | Source ↗ |
| QwQ Plus | 71% | 2026-07-16 | independent | Source ↗ |
| Qwen3-235B | 71.4% | 2026-07-16 | independent | Source ↗ |
| Grok 4 | 73.7% | 2026-07-13 | independent | Source ↗ |
| GLM-5.2 | 69.4% | 2026-07-10 | independent | Source ↗ |
| Claude Fable 5 | 75.3% | 2026-07-07 | independent | Source ↗ |
| Gemini 3 Pro | 71.8% | 2026-06-15 | independent | Source ↗ |
| o3 | 72.9% | 2026-01-17 | independent | Source ↗ |
| Mistral Large 3 | 66.3% | 2025-12-10 | independent | Source ↗ |
| Gemma 3 27B Preview | 65.5% | 2025-05-10 | independent | Source ↗ |
| Llama 4 Maverick | 70.2% | 2025-05-03 | independent | Source ↗ |
| Llama 4 Scout | 69% | 2025-05-03 | independent | Source ↗ |
| Gemini 2.5 Pro | 68.6% | 2025-04-22 | independent | Source ↗ |
| QwQ 32B | 67.2% | 2025-04-02 | independent | Source ↗ |
| GPT-4.5 | 66.3% | 2025-03-27 | independent | Source ↗ |
| Claude 3.7 Sonnet | 69.5% | 2025-03-24 | independent | Source ↗ |
| o3-preview-low (CoT + search/synthesis) | 4% | 2025-03-24 | vendor-reported | Source ↗ |
| ARChitects (Kaggle 2024 winner) | 3% | 2025-03-24 | independent | Source ↗ |
| r1 / r1-zero (single CoT) | 0.3% | 2025-03-24 | vendor-reported | Source ↗ |
| GPT-4.5 (pure LLM) | 0% | 2025-03-24 | vendor-reported | Source ↗ |
| Grok 3 | 71.8% | 2025-03-17 | independent | Source ↗ |
| Gemini 2.5 Flash (Thinking) | 66.5% | 2025-03-10 | independent | Source ↗ |
| Gemini 2.0 Flash | 64% | 2025-03-05 | independent | Source ↗ |
| Sonar Reasoning Pro | 70.2% | 2025-03-05 | independent | Source ↗ |
| Gemini 2.0 Pro | 66.7% | 2025-03-05 | independent | Source ↗ |
| o3-mini | 68.8% | 2025-02-28 | independent | Source ↗ |
| Qwen 2.5 Max | 65.9% | 2025-02-25 | independent | Source ↗ |
| Ollama DeepSeek-R1 (Q4_K_M) | 71.4% | 2025-02-19 | independent | Source ↗ |
| Kimi k1.5 | 66.3% | 2025-02-17 | independent | Source ↗ |
| Gemini 2.0 Flash Thinking | 67.9% | 2025-02-17 | independent | Source ↗ |
| DeepSeek-R1 | 68.6% | 2025-02-17 | independent | Source ↗ |
| MiniMax-Text-01 | 63.2% | 2025-02-12 | independent | Source ↗ |
| Codestral 25.01 | 57.7% | 2025-02-11 | independent | Source ↗ |
| DeepSeek-V3 | 63.9% | 2025-01-23 | independent | Source ↗ |
| Phi-4 | 59.3% | 2025-01-09 | independent | Source ↗ |
| Sonar Pro | 65.1% | 2025-01-07 | independent | Source ↗ |
| Ollama Llama 3.3 (Q4_K_M) | 64.7% | 2025-01-05 | independent | Source ↗ |
| Llama 3.3 70B | 62% | 2025-01-03 | independent | Source ↗ |
| Llama 3.3 70B Instruct | 62.8% | 2025-01-03 | independent | Source ↗ |
| o1 | 67.7% | 2025-01-02 | independent | Source ↗ |
| Amazon Nova Pro | 63.2% | 2024-12-31 | independent | Source ↗ |
| Amazon Nova Lite | 57.7% | 2024-12-31 | independent | Source ↗ |
| QwQ-32B Preview | 67.1% | 2024-12-26 | independent | Source ↗ |
| GPT-4o | 63.6% | 2024-12-18 | independent | Source ↗ |
| Sonar | 61.6% | 2024-12-13 | independent | Source ↗ |
| Qwen 2.5 Coder 32B | 60.8% | 2024-12-10 | independent | Source ↗ |
| Hunyuan-Large | 64.7% | 2024-12-03 | independent | Source ↗ |
| Claude 3.5 Sonnet (v2) | 65.1% | 2024-11-19 | independent | Source ↗ |
| Claude 3.5 Sonnet | 66.3% | 2024-11-19 | independent | Source ↗ |
| Claude 3.5 Haiku | 58.5% | 2024-11-19 | independent | Source ↗ |
| Yi-Lightning | 64% | 2024-11-12 | independent | Source ↗ |
| GLM-4-Plus | 62.4% | 2024-09-17 | independent | Source ↗ |
| Gemma 2 2B | 43.7% | 2024-08-28 | independent | Source ↗ |
| Mistral Large 2 | 62.2% | 2024-08-21 | independent | Source ↗ |
| Mistral Large 2 | 60.1% | 2024-08-21 | independent | Source ↗ |
| Llama 3.1 405B | 63.2% | 2024-08-20 | independent | Source ↗ |
| GPT-4o mini | 54.6% | 2024-08-15 | independent | Source ↗ |
| ERNIE 4.0 Turbo | 62.4% | 2024-07-26 | independent | Source ↗ |
| Gemma 2 9B | 55.4% | 2024-07-25 | independent | Source ↗ |
| Gemma 2 27B | 60.1% | 2024-07-25 | independent | Source ↗ |
| GPT-5.6 Sol (Max) | 92.5% | 2024-06-01 | vendor-reported | Source ↗ |
| GPT-5.5 (xHigh) | 85% | 2024-06-01 | vendor-reported | Source ↗ |
| Claude 4.7 (Max) | 75.8% | 2024-06-01 | vendor-reported | Source ↗ |
| NVARC | 24% | 2024-06-01 | independent | Source ↗ |
| Snowflake Arctic | 58.5% | 2024-05-22 | independent | Source ↗ |
| DBRX Instruct | 59.3% | 2024-04-24 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human solve rate on updated 2025 ARC-AGI visual logic puzzles with increased transformation complexity.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.task success rate on the public evaluation split (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks arc-agi-2 --batch_size autoopencompass --datasets arc-agi-2 --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for ARC-AGI-2, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace