SimpleQA
4,326 short fact-seeking questions chosen to stump GPT-4 โ a hallucination test graded correct / incorrect / not-attempted.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
4,326 short fact-seeking questions chosen to stump GPT-4 โ a hallucination test graded correct / incorrect / not-attempted.
~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models โ designed to be the final closed-ended academic benchmark.
MMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.
PhD-written science questions so hard that skilled non-experts with Google score 34% โ the 'Google-proof' exam, reported on its 198-question Diamond subset.
57-subject multiple-choice exam spanning STEM, humanities, and social sciences โ the defining knowledge benchmark of the 2020โ2024 era.
A manual error audit of MMLU: experts re-annotated thousands of questions and found ~6.5% are broken โ the evidence base for MMLU's label-noise ceiling.