Soohak
Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.
Original, unpublished research-level math problems by expert mathematicians โ mostly private to resist contamination. Frontier models solved under 2% at launch.
A continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.
Formal theorem-proving on Putnam problems โ models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.
Bilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.
The MAA's competition exam โ 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.
A human-translated 250-problem GSM8K subset for comparing grade-school mathematical reasoning across ten languages.
Grade-level math word problems that require combining a question with structured, textual, or image-rendered tables.
8.5K grade-school math word problems (2โ8 arithmetic steps) โ the default elementary-math benchmark and the classic chain-of-thought demonstration.
12,500 AMC/AIME-level competition problems with worked solutions โ usually evaluated on its 500-problem MATH-500 subset, the reasoning-model math standard.
A 4,428-problem Olympiad mathematics suite spanning more than 33 subdomains and ten annotated difficulty levels.
GSM8K test problems with unusually large numerical values, designed to separate reasoning from easy arithmetic.