Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 16 of 16 benchmarks (filtered from 63)
ActiveCoding

FrontierCode

Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.

Trust Score
95
Contaminationlow
ActiveCoding

SWE-bench Pro

Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.

Trust Score
90
Contaminationmedium
ActiveCoding

Terminal-Bench

Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.

Trust Score
80
Contaminationhigh
Nearing SaturationCoding

MBPP+

A cleaned 378-task MBPP evaluation with roughly 35 times more tests for stricter Python functional correctness.

Trust Score
65
Contaminationhigh
SaturatedCoding

LiveCodeBench

Continuously updated contest problems for time-windowed code generation, execution, test prediction, and self-repair evaluation.

Trust Score
35
Contaminationmedium
SaturatedCoding

MBPP

Mostly Basic Programming Problems: 974 crowd-sourced Python synthesis tasks intended for entry-level programmers.

Trust Score
25
Contaminationhigh
SaturatedCoding

SWE-bench Verified

A 500-task, engineer-reviewed subset of SWE-bench that became the standard repository-level coding-agent test.

Trust Score
25
Contaminationhigh
SaturatedCoding

HumanEval

OpenAI's 164 hand-written Python function-synthesis tasks, graded by executing generated completions against unit tests.

Trust Score
25
Contaminationhigh
SaturatedCoding

HumanEval+

HumanEval's 164 Python tasks re-evaluated with roughly 80 times more tests to expose plausible but incorrect programs.

Trust Score
25
Contaminationhigh
DeprecatedCoding

SWE-Lancer

Current 198-task offline patch benchmark derived from paid Upwork work, with historical mixed implementation and manager-task variants.

Trust Score
20
Contaminationmedium
DeprecatedCoding

CodeXGLUE

A broad code-intelligence suite spanning understanding, retrieval, completion, translation, repair, generation, and summarization.

Trust Score
20
Contaminationmedium
DeprecatedCoding

Mercury

Python code-synthesis benchmark that rewards both functional correctness and runtime efficiency against real solution distributions.

Trust Score
10
Contaminationhigh
DeprecatedCoding

MultiPL-E

Parallel HumanEval and MBPP translations for comparing code generation across many programming languages.

Trust Score
10
Contaminationhigh
DeprecatedCoding

SWE-bench

Real GitHub issues from 12 Python repositories, solved by generating patches that must pass repository tests.

Trust Score
10
Contaminationhigh
DeprecatedCoding

CodeContests

DeepMind's competitive-programming corpus with temporally split problems, human submissions, and generated correctness tests.

Trust Score
10
Contaminationhigh
DeprecatedCoding

CyberSecEval

Meta's evolving suite for insecure code, cyberattack assistance, prompt injection, exploitation, SOC analysis, and autopatching.

Trust Score
10
Contaminationhigh