HellaSwag
Choose the plausible continuation of an everyday event from one real ending and three adversarially filtered machine generations.
GPT-4 base reached 95.3% on a private holdout against the original 95.6% human reference; the official leaderboard closed submissions in 2024, and public validation data is contamination-prone.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| GPT-4o | 95.3% | 2024-08-01 | vendor-reported | Source ↗ |
| Llama 3 70B | 88% | 2024-04-18 | vendor-reported | Source ↗ |
| Claude 3 Opus | 95.4% | 2024-03-04 | vendor-reported | Source ↗ |
| Gemini Ultra (10-shot decontaminated) | 87.8% | 2023-12-06 | vendor-reported | Source ↗ |
| GPT-4 base (10-shot) | 95.3% | 2023-03-14 | vendor-reported | Source ↗ |
| LLaMA-65B | 84.2% | 2023-02-27 | vendor-reported | Source ↗ |
| PaLM 540B | 83.4% | 2022-04-01 | vendor-reported | Source ↗ |
| GPT-3 175B (few-shot) | 79.3% | 2020-07-22 | vendor-reported | Source ↗ |
| RoBERTa | 85.2% | 2019-07-25 | independent | Source ↗ |
| BERT-Large | 47.3% | 2019-05-19 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human accuracy on adversarial sentence completion and commonsense story continuation.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.normalized accuracy (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks hellaswag --batch_size autoopencompass --datasets hellaswag --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for HellaSwag, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace