AIME, Olympiad Proofs & Formal Mathematics

Mathematics & Olympiad Problem Solving

Mathematical reasoning has advanced from elementary word problems to solving American Invitational Mathematics Examination (AIME) and Putnam-level research problems through extended test-time verification.

Human Baseline40.0% (USAMO Qualifier Baseline)
Contamination StatusTotal on GSM8K/MATH-500; Low on FrontierMath & AIME 2026
Top Frontier ModelGPT-5.6 Sol (96.2%) & Claude Fable 5 (96.0%)

Forensic Analysis & Methodology

Why Classic Benchmarks Failed

GSM8K reached 97%+ saturation across virtually all frontier models due to verbatim training data inclusion and low problem complexity.

The TrustTheBench & LiveBench Approach

FrontierMath and newly released AIME 2025/2026 exams test rigorous multi-step numerical calculation and symbolic deduction with near-zero pre-training contamination.

Verified Category Leaderboard

Ranked strictly by verified Mathematics & Olympiad Problem Solving performance.

RankModelResearch LabAccess TypeCategory ScoreSpeedPricing / 1M
#1Claude Fable 5Anthropicproprietary96.0%45 tok/s$10.00 in
#2GPT-5.5 ThinkingOpenAIproprietary95.9%70 tok/s$3.00 in
#3Claude Opus 5Anthropicproprietary95.7%48 tok/s$5.00 in
#4Grok 4xAIproprietary95.0%75 tok/s$5.00 in
#5GPT-5OpenAIproprietary94.0%85 tok/s$2.50 in
#6Gemini 3.7 FlashGoogle DeepMindproprietary93.5%182 tok/s$0.75 in
#7QwQ PlusAlibaba Cloud / Qwenopen-weight92.5%52 tok/s$0.80 in
#8Grok 4.6xAIproprietary92.1%75 tok/s$2.00 in
#9Gemini 3 ProGoogle DeepMindproprietary92.0%80 tok/s$2.50 in
#10Claude Sonnet 5Anthropicproprietary91.5%95 tok/s$3.00 in
#11DeepSeek V4-ProDeepSeekopen-weight91.4%65 tok/s$0.66 in
#12Gemini 3.6 FlashGoogle DeepMindproprietary91.2%175 tok/s$0.75 in
#13Qwen3-235BAlibaba Cloud / Qwenopen-weight91.0%70 tok/s$0.80 in
#14GLM-5.2Zhipu AIproprietary90.5%72 tok/s$1.50 in
#15GPT-6 Astra (max)OpenAIproprietary90.2%62 tok/s$4.50 in

Deployment Recommendations

🏆 Best Overall Frontier

GPT-5.6 Sol (96.2%) & Claude Fable 5 (96.0%)

Maximum reasoning depth and lowest error rate on difficult boundary problems.

🔓 Best Open-Weight

QwQ Plus (92.5%) & DeepSeek-R1 (90.8%)

Host locally or via cost-effective cloud providers with full data privacy.

⚡ Best Value & Speed

Gemini 3.7 Flash (93.5% at $0.75/1M)

Optimal price-to-performance ratio for high-volume automated agent pipelines.

Relevant Benchmark References

AIME 2026

Metric: Solve Rate %View Methodology →

FrontierMath

Metric: Verified Solve %View Methodology →

LiveBench Math

Metric: Exact Match %View Methodology →