CyberSecEval
Meta's evolving suite for insecure code, cyberattack assistance, prompt injection, exploitation, SOC analysis, and autopatching.
The original CyberSecEval benchmark is no longer the recommended standalone version: Meta advanced through CyberSecEval 2 and 3 to CyberSecEval 4, changing tasks and metrics rather than maintaining one comparable score axis.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| llama 3 70b-instruct | 29.27% | 2024-04-18 | vendor-reported | Source ↗ |
| llama 3 8b-instruct | 45.27% | 2024-04-18 | vendor-reported | Source ↗ |
| UN codellama-70b-instruct | 12.93% | 2024-01-29 | vendor-reported | Source ↗ |
| UN codellama-34b-instruct | 36.33% | 2023-08-24 | vendor-reported | Source ↗ |
| UN codellama-13b-instruct | 37.27% | 2023-08-24 | vendor-reported | Source ↗ |
| GPT-4 | 19.87% | 2023-03-14 | vendor-reported | Source ↗ |
| gpt-3.5-turbo | 39.13% | 2023-03-01 | vendor-reported | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human cyber-security specialist vulnerability detection and exploit identification rate.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.average textual prompt-injection success rate (%) — lower is saferDataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks cyberseceval --batch_size autoopencompass --datasets cyberseceval --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for CyberSecEval, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace