FrontierMath
Original, unpublished research-level math problems by expert mathematicians — mostly private to resist contamination. Frontier models solved under 2% at launch.
The strongest active case on the wiki. At launch (Nov 2024) state-of-the-art models solved UNDER 2% — a benchmark frontier systems almost entirely fail. Reasoning models with Python tools have since climbed (OpenAI announced o3 at 25.2% in Dec 2024, tools-on), but that figure is contested — Epoch's own independent reproductions of comparable models came in far lower (~11%). By the June 2026 v2 release Epoch added a harder Tier 4 precisely because Tiers 1–3 progress had become significant, yet the hardest tier remains far from solved. Current figures move fast and are tool/compute-dependent — consult Epoch's live dashboard; this page does not restate unverified third-party numbers. Saturation here would mean research-level, closed-form mathematics is largely automated.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Elite mathematics competition winners and PhD researchers score <2% without computing tools, ~25% with Python.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy — % of problems solved, automated answer verification (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks frontiermath --batch_size autoopencompass --datasets frontiermath --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for FrontierMath, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace