MathematicsNearing Saturation
Trust Score:
75

AIME

The MAA's competition exam — 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.

Launched: Refresh: yearly
Status Assessment (nearing-saturation):

AIME's status is genuinely split by its yearly refresh. OLD exams are saturated and contaminated: reasoning models solved AIME 2024 (o3 ~96.7%, up from GPT-4o's 13.4%). Each FRESH exam has a shrinking clean window, but frontier models now score high even on fresh exams (AIME 2025 ~85–91%). So the rolling benchmark is nearing-saturation — no single exam discriminates for long, yet the yearly refresh keeps briefly-clean signal alive. The plotted gauge/timeline is the AIME 2024 progression (one exam); the perpetual benchmark's yearly reset is described in prose, since the signature gauge shows one monotonic climb and cannot depict a reset.

Performance Timeline

Longitudinal progression of model scores against human baselines.

Performance & Historical Trajectory

Empirical score progression across model release dates and evaluation rounds.

Independent Vendor-Reported Human Baseline (13.3%)
ModelScoreDateSource TypeProvenance
o3 (AIME 2024, pass@1)96.7%2024-12-20vendor-reportedSource ↗
GPT-4o (AIME 2024, pass@1)13.4%2024-09-12vendor-reportedSource ↗
o1 (AIME 2024, pass@1)74.4%2024-09-12vendor-reportedSource ↗

Human Baseline & Difficulty Horizon

Calibrated human reference points, specialist benchmarks, and ceiling thresholds.
Measured Human Score13.3%Domain Expert Baseline
Baseline Protocol & Interpretation

Average score among the top 2.5% US high school mathematics competition qualifiers.

Metric & Scoring Methodology

Verification protocols, aggregation formulas, and specialized metric variants.
Primary Metric:accuracy — every answer is an integer 000–999, exact match (%)
Scoring Engine:exact-match

Dataset & Compute Cost

Evaluation volume, public availability, API pricing, and local hardware requirements.
Total Dataset Size30Annotated evaluation items
Public Test SetPublicOpenly mirrored on repositories
Access GatingOpen AccessUnrestricted download
Evaluation LicenseOpen AccessDataset usage and redistribution terms
Frontier API Compute Cost:

$5 – $20 USD for full benchmark evaluation run on frontier APIs.

Recommended Local GPU Setup:

1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang

Official Dataset & Benchmark Files:Download / View Dataset Repository ↗

How to Run & Reproduce

Standardized evaluation protocols, CLI commands, and reproducible runner templates.
Prompt Regimezero-shot
Reasoning Modedirect
Sampling Temp0
Pass@k Budgetk = 1
Tools & SandboxPure Text
Scoring Verifierexact-match
Option AEleutherAI LM-Evaluation-Harness (Open-Weight Models)
lm_eval --model hf --model_args pretrained=<model_path> --tasks aime --batch_size auto
Option BOpenCompass Evaluation Framework
opencompass --datasets aime --models <model_config>
Python APIDeterministic Inference Loop Snippet
# Standard API Evaluation Loop
from openai import OpenAI

client = OpenAI()
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": prompt}],
    temperature=0.0,
)
Standardized Reporting Requirement:

When publishing results for AIME, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.

Contamination & Memorization Analysis

Audit of pretraining exposure risks, memorization vectors, and refresh policies.
Overall Contamination Risk:MEDIUM
Refresh Cadence:

Static fixed snapshot

Test Set Exposure:

Public on web / HuggingFace