Soohak
Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.
Evaluates instruction following for enterprise/API pipelines โ exact output format, strict step ordering, ranking, and calibrated refusal across six categories.
Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.
Original, unpublished research-level math problems by expert mathematicians โ mostly private to resist contamination. Frontier models solved under 2% at launch.
A continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.
Formal theorem-proving on Putnam problems โ models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.
4,326 short fact-seeking questions chosen to stump GPT-4 โ a hallucination test graded correct / incorrect / not-attempted.
Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.
Interactive ARC environments that score agents by human-relative action efficiency, planning, exploration, and adaptation.
~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models โ designed to be the final closed-ended academic benchmark.
Bilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.
MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.
Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.
Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.
Bilingual benchmark of how models follow instructions when several constraints must hold at once โ where stacking constraints is not linear.
707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints โ whether instruction following survives deployment-scale prompts.
Realistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.
Multiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.
MMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.
80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1โ10 scale.
155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).
Zero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.
The MAA's competition exam โ 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.
The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.