Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 11 of 11 benchmarks (filtered from 63)
ActiveLong Context

OpenAI-MRCR

MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.

Trust Score
80
Contaminationhigh
Nearing SaturationLong Context

RULER

Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.

Trust Score
80
Contaminationlow
ActiveLong Context

LongBench-Pro

Realistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.

Trust Score
80
Contaminationhigh
Nearing SaturationLong Context

QuALITY

Multiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.

Trust Score
75
Contaminationmedium
Nearing SaturationLong Context

ZeroSCROLLS

Zero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.

Trust Score
75
Contaminationmedium
Nearing SaturationLong Context

QASPER

Scientific-paper QA over full NLP papers, with abstractive, extractive, yes/no, and unanswerable answers plus supporting evidence.

Trust Score
65
Contaminationhigh
Nearing SaturationLong Context

QMSum

Query-focused meeting summarization over long transcripts from academic, product, and committee meetings.

Trust Score
65
Contaminationhigh
Nearing SaturationLong Context

SCROLLS

Seven-task long-sequence suite covering summarization, QA, and NLI over naturally long English texts.

Trust Score
65
Contaminationhigh
Nearing SaturationLong Context

InfiniteBench

100k+ token benchmark mixing realistic and synthetic tasks across English, Chinese, code, math, and retrieval.

Trust Score
65
Contaminationhigh
Nearing SaturationLong Context

LongBench

Bilingual multitask suite for long-context understanding across QA, summarization, few-shot learning, synthetic tasks, and code completion.

Trust Score
65
Contaminationhigh
SaturatedLong Context

Needle-in-a-Haystack

Synthetic retrieval stress test that hides facts at different depths and context lengths, then asks the model to recover them.

Trust Score
40
Contaminationlow