Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
Displaying 24 of 63 benchmarks
ActiveMathematics

Soohak

Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.

Trust Score
95
Contaminationlow
ActiveInstruction Following

FireBench

Evaluates instruction following for enterprise/API pipelines โ€” exact output format, strict step ordering, ranking, and calibrated refusal across six categories.

Trust Score
95
Contaminationlow
ActiveCoding

FrontierCode

Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.

Trust Score
95
Contaminationlow
ActiveMathematics

FrontierMath

Original, unpublished research-level math problems by expert mathematicians โ€” mostly private to resist contamination. Frontier models solved under 2% at launch.

Trust Score
95
Contaminationlow
ActiveMathematics

MathArena

A continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.

Trust Score
95
Contaminationlow
ActiveMathematics

PutnamBench

Formal theorem-proving on Putnam problems โ€” models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.

Trust Score
90
Contaminationmedium
ActiveKnowledge

SimpleQA

4,326 short fact-seeking questions chosen to stump GPT-4 โ€” a hallucination test graded correct / incorrect / not-attempted.

Trust Score
90
Contaminationmedium
ActiveCoding

SWE-bench Pro

Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.

Trust Score
90
Contaminationmedium
ActiveAgentic & Tool Use

ARC-AGI-3

Interactive ARC environments that score agents by human-relative action efficiency, planning, exploration, and adaptation.

Trust Score
90
Contaminationmedium
ActiveKnowledge

Humanity's Last Exam

~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models โ€” designed to be the final closed-ended academic benchmark.

Trust Score
90
Contaminationmedium
ActiveMathematics

OlympiadBench

Bilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.

Trust Score
80
Contaminationhigh
ActiveLong Context

OpenAI-MRCR

MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.

Trust Score
80
Contaminationhigh
Nearing SaturationLong Context

RULER

Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.

Trust Score
80
Contaminationlow
ActiveCoding

Terminal-Bench

Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.

Trust Score
80
Contaminationhigh
Nearing SaturationInstruction Following

ComplexBench

Bilingual benchmark of how models follow instructions when several constraints must hold at once โ€” where stacking constraints is not linear.

Trust Score
80
Contaminationlow
Nearing SaturationInstruction Following

AgentIF

707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints โ€” whether instruction following survives deployment-scale prompts.

Trust Score
80
Contaminationlow
ActiveLong Context

LongBench-Pro

Realistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.

Trust Score
80
Contaminationhigh
Nearing SaturationLong Context

QuALITY

Multiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.

Trust Score
75
Contaminationmedium
Nearing SaturationKnowledge

MMLU-Pro

MMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.

Trust Score
75
Contaminationmedium
Nearing SaturationInstruction Following

MT-Bench

80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1โ€“10 scale.

Trust Score
75
Contaminationmedium
Nearing SaturationInstruction Following

StructFlowBench

155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).

Trust Score
75
Contaminationmedium
Nearing SaturationLong Context

ZeroSCROLLS

Zero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.

Trust Score
75
Contaminationmedium
Nearing SaturationMathematics

AIME

The MAA's competition exam โ€” 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.

Trust Score
75
Contaminationmedium
Nearing SaturationReasoning

ARC-AGI-2

The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.

Trust Score
75
Contaminationmedium