Mathematics & Olympiad Problem Solving
Mathematical reasoning has advanced from elementary word problems to solving American Invitational Mathematics Examination (AIME) and Putnam-level research problems through extended test-time verification.
Forensic Analysis & Methodology
Why Classic Benchmarks Failed
GSM8K reached 97%+ saturation across virtually all frontier models due to verbatim training data inclusion and low problem complexity.
The TrustTheBench & LiveBench Approach
FrontierMath and newly released AIME 2025/2026 exams test rigorous multi-step numerical calculation and symbolic deduction with near-zero pre-training contamination.
Verified Category Leaderboard
Ranked strictly by verified Mathematics & Olympiad Problem Solving performance.
| Rank | Model | Research Lab | Access Type | Category Score | Speed | Pricing / 1M |
|---|---|---|---|---|---|---|
| #1 | Claude Fable 5 | Anthropic | proprietary | 96.0% | 45 tok/s | $10.00 in |
| #2 | GPT-5.5 Thinking | OpenAI | proprietary | 95.9% | 70 tok/s | $3.00 in |
| #3 | Claude Opus 5 | Anthropic | proprietary | 95.7% | 48 tok/s | $5.00 in |
| #4 | Grok 4 | xAI | proprietary | 95.0% | 75 tok/s | $5.00 in |
| #5 | GPT-5 | OpenAI | proprietary | 94.0% | 85 tok/s | $2.50 in |
| #6 | Gemini 3.7 Flash | Google DeepMind | proprietary | 93.5% | 182 tok/s | $0.75 in |
| #7 | QwQ Plus | Alibaba Cloud / Qwen | open-weight | 92.5% | 52 tok/s | $0.80 in |
| #8 | Grok 4.6 | xAI | proprietary | 92.1% | 75 tok/s | $2.00 in |
| #9 | Gemini 3 Pro | Google DeepMind | proprietary | 92.0% | 80 tok/s | $2.50 in |
| #10 | Claude Sonnet 5 | Anthropic | proprietary | 91.5% | 95 tok/s | $3.00 in |
| #11 | DeepSeek V4-Pro | DeepSeek | open-weight | 91.4% | 65 tok/s | $0.66 in |
| #12 | Gemini 3.6 Flash | Google DeepMind | proprietary | 91.2% | 175 tok/s | $0.75 in |
| #13 | Qwen3-235B | Alibaba Cloud / Qwen | open-weight | 91.0% | 70 tok/s | $0.80 in |
| #14 | GLM-5.2 | Zhipu AI | proprietary | 90.5% | 72 tok/s | $1.50 in |
| #15 | GPT-6 Astra (max) | OpenAI | proprietary | 90.2% | 62 tok/s | $4.50 in |
Deployment Recommendations
GPT-5.6 Sol (96.2%) & Claude Fable 5 (96.0%)
Maximum reasoning depth and lowest error rate on difficult boundary problems.
QwQ Plus (92.5%) & DeepSeek-R1 (90.8%)
Host locally or via cost-effective cloud providers with full data privacy.
Gemini 3.7 Flash (93.5% at $0.75/1M)
Optimal price-to-performance ratio for high-volume automated agent pipelines.