Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 6 of 6 benchmarks (filtered from 63)
ActiveInstruction Following

FireBench

Evaluates instruction following for enterprise/API pipelines โ€” exact output format, strict step ordering, ranking, and calibrated refusal across six categories.

Trust Score
95
Contaminationlow
Nearing SaturationInstruction Following

ComplexBench

Bilingual benchmark of how models follow instructions when several constraints must hold at once โ€” where stacking constraints is not linear.

Trust Score
80
Contaminationlow
Nearing SaturationInstruction Following

AgentIF

707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints โ€” whether instruction following survives deployment-scale prompts.

Trust Score
80
Contaminationlow
Nearing SaturationInstruction Following

MT-Bench

80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1โ€“10 scale.

Trust Score
75
Contaminationmedium
Nearing SaturationInstruction Following

StructFlowBench

155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).

Trust Score
75
Contaminationmedium
DeprecatedInstruction Following

IFEval

~500 prompts, each carrying an explicit, machine-checkable constraint (format, content, or style) that a response must respect.

Trust Score
25
Contaminationlow