MBPP
Mostly Basic Programming Problems: 974 crowd-sourced Python synthesis tasks intended for entry-level programmers.
CodeSOTA's 2026 MBPP pass@1 registry has multiple frontier and open-code models clustered from 89% to 95%, while the short public tasks and three-test protocol remain too weak for current frontier comparison.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| o3-mini (CodeSOTA MBPP pass@1) | 93.3% | 2025-01-31 | independent | Source ↗ |
| DeepSeek-V3 (CodeSOTA MBPP pass@1) | 89.3% | 2024-12-26 | independent | Source ↗ |
| Qwen2.5-Coder-32B-Instruct (CodeSOTA MBPP pass@1) | 90.2% | 2024-11-12 | independent | Source ↗ |
| Qwen2.5-Coder 32B (CodeSOTA MBPP pass@1) | 90.2% | 2024-11-12 | independent | Source ↗ |
| Claude 3.5 Sonnet (Oct 2024; CodeSOTA MBPP pass@1) | 91% | 2024-10-22 | independent | Source ↗ |
| DeepSeek-Coder-V2-Instruct (CodeSOTA MBPP pass@1) | 89.4% | 2024-06-14 | independent | Source ↗ |
| o4-mini (CodeSOTA MBPP pass@1) | 94.9% | 2024-06-01 | independent | Source ↗ |
| Claude Opus 4 (CodeSOTA MBPP pass@1) | 92% | 2024-06-01 | independent | Source ↗ |
| GPT-4.1 (CodeSOTA MBPP pass@1) | 90.9% | 2024-06-01 | independent | Source ↗ |
| Claude Sonnet 4 (CodeSOTA MBPP pass@1) | 89.6% | 2024-06-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human software engineer pass@1 solve rate on beginner-to-intermediate Python tasks without execution feedback.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.pass@1 on the 500-task test split (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks mbpp --batch_size autoopencompass --datasets mbpp --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for MBPP, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace