Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 12 of 12 benchmarks (filtered from 63)
ActiveMathematics

Soohak

Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.

Trust Score
95
Contaminationlow
ActiveMathematics

FrontierMath

Original, unpublished research-level math problems by expert mathematicians โ€” mostly private to resist contamination. Frontier models solved under 2% at launch.

Trust Score
95
Contaminationlow
ActiveMathematics

MathArena

A continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.

Trust Score
95
Contaminationlow
ActiveMathematics

PutnamBench

Formal theorem-proving on Putnam problems โ€” models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.

Trust Score
90
Contaminationmedium
ActiveMathematics

OlympiadBench

Bilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.

Trust Score
80
Contaminationhigh
Nearing SaturationMathematics

AIME

The MAA's competition exam โ€” 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.

Trust Score
75
Contaminationmedium
Nearing SaturationMathematics

MGSM

A human-translated 250-problem GSM8K subset for comparing grade-school mathematical reasoning across ten languages.

Trust Score
65
Contaminationhigh
SaturatedMathematics

TabMWP

Grade-level math word problems that require combining a question with structured, textual, or image-rendered tables.

Trust Score
25
Contaminationhigh
SaturatedMathematics

GSM8K

8.5K grade-school math word problems (2โ€“8 arithmetic steps) โ€” the default elementary-math benchmark and the classic chain-of-thought demonstration.

Trust Score
25
Contaminationhigh
SaturatedMathematics

MATH

12,500 AMC/AIME-level competition problems with worked solutions โ€” usually evaluated on its 500-problem MATH-500 subset, the reasoning-model math standard.

Trust Score
25
Contaminationhigh
DeprecatedMathematics

Omni-MATH

A 4,428-problem Olympiad mathematics suite spanning more than 33 subdomains and ten annotated difficulty levels.

Trust Score
10
Contaminationhigh
DeprecatedMathematics

GSM-Hard

GSM8K test problems with unusually large numerical values, designed to separate reasoning from easy arithmetic.

Trust Score
10
Contaminationhigh