ARC-AGI
Colored-grid abstraction tasks that test few-shot rule induction, not memorized facts, language fluency, or school math.
ARC-AGI-1 remains useful as a historical abstraction benchmark, but the official ARC Prize leaderboard generated on 2026-07-23 lists multiple verified systems above 95% on the ARC-AGI-1 semi-private axis, including GPT-5.6 Sol at 97.5%. ARC Prize launched ARC-AGI-2 because ARC-AGI-1 had become too brute-forceable and less discriminative for frontier systems.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| o3-preview-low (CoT + search/synthesis) | 75.7% | 2024-12-20 | vendor-reported | Source ↗ |
| Jeremy Berman system | 53.6% | 2024-12-18 | independent | Source ↗ |
| o1-preview | 18% | 2024-09-13 | vendor-reported | Source ↗ |
| GPT-4o | 5% | 2024-06-27 | vendor-reported | Source ↗ |
| Claude 4.7 (High) | 93.5% | 2024-06-01 | vendor-reported | Source ↗ |
| GPT-5.5 Pro (High) | 96.5% | 2024-06-01 | vendor-reported | Source ↗ |
| GPT-5.6 Sol (xHigh) | 97.5% | 2024-06-01 | vendor-reported | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Standard human solve rate on abstract visual reasoning grid transformation puzzles without training.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.task success rate on the evaluation split (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks arc-agi --batch_size autoopencompass --datasets arc-agi --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for ARC-AGI, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace