Terminal-Bench
Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.
Terminal-Bench 2.1 remains actively maintained and submission-driven; its verified 17-entry leaderboard spans 58.7% to 83.8% across multiple agent-model pairs, while Terminal-Bench 3.0 is still in development.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| Codex + GPT-5.6 Luna (max; TB 2.1) | 75.7% | 2024-06-01 | independent | Source ↗ |
| Claude Code + GLM-5.1 (max; TB 2.1) | 58.7% | 2024-06-01 | independent | Source ↗ |
| Terminus 2 + Gemini 3.1 Pro (high; TB 2.1) | 65.6% | 2024-06-01 | independent | Source ↗ |
| Gemini CLI + Gemini 3.1 Pro (high; TB 2.1) | 65.8% | 2024-06-01 | independent | Source ↗ |
| Gemini CLI + Gemini 3 Pro (high; TB 2.1) | 65.8% | 2024-06-01 | independent | Source ↗ |
| Terminus 2 + Opus 4.7 (max; TB 2.1) | 66.1% | 2024-06-01 | independent | Source ↗ |
| Claude Code + Opus 4.7 (max; TB 2.1) | 68.9% | 2024-06-01 | independent | Source ↗ |
| Terminus 2 + Gemini 3 Pro (high; TB 2.1) | 73.9% | 2024-06-01 | independent | Source ↗ |
| Claude Code + Sonnet 5 (high; TB 2.1) | 74.6% | 2024-06-01 | independent | Source ↗ |
| Claude Code + Fable 5 (xhigh; TB 2.1) | 83.8% | 2024-06-01 | independent | Source ↗ |
| mini-SWE-agent + Muse Spark 1.1 (xhigh; TB 2.1) | 76.2% | 2024-06-01 | independent | Source ↗ |
| Terminus 2 + GPT-5.5 (xhigh; TB 2.1) | 78% | 2024-06-01 | independent | Source ↗ |
| Codex + GPT-5.6 Terra (max; TB 2.1) | 78.4% | 2024-06-01 | independent | Source ↗ |
| Claude Code + Opus 4.8 (high; TB 2.1) | 78.9% | 2024-06-01 | independent | Source ↗ |
| Cursor CLI + Grok 4.5 (high; TB 2.1) | 79.3% | 2024-06-01 | independent | Source ↗ |
| Terminus 2 + Fable 5 (high; TB 2.1) | 80.4% | 2024-06-01 | independent | Source ↗ |
| Codex + GPT-5.5 (xhigh; TB 2.1) | 83.1% | 2024-06-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human systems engineer success rate executing complex multi-step terminal and shell commands.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.task accuracy on Terminal-Bench 2.1 (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks terminal-bench --batch_size autoopencompass --datasets terminal-bench --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for Terminal-Bench, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace