Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 10 of 10 benchmarks (filtered from 63)
Nearing SaturationReasoning

ARC-AGI-2

The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.

Trust Score
75
Contaminationmedium
SaturatedReasoning

SuperGLUE

Eight English NLU tasks spanning question answering, inference, causal reasoning, word sense, and coreference.

Trust Score
25
Contaminationhigh
SaturatedReasoning

WinoGrande

Binary fill-in-the-blank commonsense problems, scaled from Winograd schemas and adversarially filtered to reduce dataset shortcuts.

Trust Score
25
Contaminationhigh
SaturatedReasoning

GLUE

Nine-task English NLU suite combining acceptability, sentiment, similarity, paraphrase, inference and coreference into one score.

Trust Score
25
Contaminationhigh
SaturatedReasoning

AGIEval

Standardized human exams โ€” SATs, LSATs, China's Gaokao and civil-service tests โ€” repurposed to grade models against real human test-takers.

Trust Score
25
Contaminationhigh
SaturatedReasoning

ARC-AGI

Colored-grid abstraction tasks that test few-shot rule induction, not memorized facts, language fluency, or school math.

Trust Score
25
Contaminationhigh
SaturatedReasoning

BIG-Bench Hard

23 BIG-Bench tasks where models trailed humans โ€” the benchmark that proved chain-of-thought prompting works, then fell to the reasoning-model era.

Trust Score
25
Contaminationhigh
SaturatedReasoning

CommonsenseQA

Five-choice questions derived from ConceptNet relations, designed to require everyday knowledge beyond a supplied passage.

Trust Score
25
Contaminationhigh
SaturatedReasoning

HellaSwag

Choose the plausible continuation of an everyday event from one real ending and three adversarially filtered machine generations.

Trust Score
25
Contaminationhigh
DeprecatedReasoning

BIG-Bench

204-task community-built mega-suite spanning reasoning, knowledge, math, code and social bias โ€” the sprawling proving ground BBH was distilled from.

Trust Score
10
Contaminationhigh