Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 1 of 1 benchmarks (filtered from 63)
ActiveAgentic & Tool Use

ARC-AGI-3

Interactive ARC environments that score agents by human-relative action efficiency, planning, exploration, and adaptation.

Trust Score
90
Contaminationmedium