LiveCodeBench
Continuously updated contest problems for time-windowed code generation, execution, test prediction, and self-repair evaluation.
In the August 2, 2026 LiveCodeBench leaderboard snapshot, Qwen3.7 Max reaches 91.6%, Qwen3.7 Plus 89.6%, and GLM-4.7 84.9%; the leading cluster leaves too little reliable headroom for frontier ranking on this snapshot.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| DeepSeek V3 (open weight) | 37.6% | 2024-12-26 | independent | Source ↗ |
| Qwen3.7 Max (qwen3-7-max; closed) | 91.6% | 2024-06-01 | independent | Source ↗ |
| Qwen3.7 Plus (qwen3-7-plus; closed) | 89.6% | 2024-06-01 | independent | Source ↗ |
| GLM-4.7 (glm-4-7; open weight) | 84.9% | 2024-06-01 | independent | Source ↗ |
| Qwen3.6-27B (open weight) | 83.9% | 2024-06-01 | independent | Source ↗ |
| Qwen3.6-35B-A3B (open weight) | 80.4% | 2024-06-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Average solve rate of rated human participants across live LeetCode, AtCoder, and Codeforces contests.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.pass@1 on code generation (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks livecodebench --batch_size autoopencompass --datasets livecodebench --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for LiveCodeBench, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace