Humanity's Last Exam
~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models — designed to be the final closed-ended academic benchmark.
The clearest active benchmark on the wiki. Launch scores were single digits (GPT-4o 3.3%, o1 9.1%, DeepSeek-R1 9.4%, no tools); the current no-tools frontier is ~38% (Gemini 3 Pro, Nov 2025) — a steep climb but a long way from solved. Given the 'last exam' thesis, saturation here would mean closed-ended academic testing has run out of headroom entirely, so 'what replaces it' points toward open-ended, agentic, and interactive evaluation rather than a harder Q&A set. Note: tool-augmented and agent scores run much higher and are tracked separately (see Timeline).
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| GPT-4o | 3.3% | 2025-01-24 | independent | Source ↗ |
| Claude 3.5 Sonnet | 4.3% | 2025-01-24 | independent | Source ↗ |
| o1 | 9.1% | 2025-01-24 | independent | Source ↗ |
| DeepSeek-R1 | 9.4% | 2025-01-24 | independent | Source ↗ |
| Grok 4 | 24.5% | 2024-06-01 | independent | Source ↗ |
| GPT-5 | 25.3% | 2024-06-01 | independent | Source ↗ |
| Gemini 3 Pro | 38.3% | 2024-06-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Cross-disciplinary academic experts score ~84% across questions outside their exact narrow specialty (authors achieve ~95%).
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy, no tools (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks humanitys-last-exam --batch_size autoopencompass --datasets humanitys-last-exam --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for Humanity's Last Exam, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace