AIME
The MAA's competition exam — 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.
AIME's status is genuinely split by its yearly refresh. OLD exams are saturated and contaminated: reasoning models solved AIME 2024 (o3 ~96.7%, up from GPT-4o's 13.4%). Each FRESH exam has a shrinking clean window, but frontier models now score high even on fresh exams (AIME 2025 ~85–91%). So the rolling benchmark is nearing-saturation — no single exam discriminates for long, yet the yearly refresh keeps briefly-clean signal alive. The plotted gauge/timeline is the AIME 2024 progression (one exam); the perpetual benchmark's yearly reset is described in prose, since the signature gauge shows one monotonic climb and cannot depict a reset.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Average score among the top 2.5% US high school mathematics competition qualifiers.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy — every answer is an integer 000–999, exact match (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks aime --batch_size autoopencompass --datasets aime --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for AIME, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace