FireBench
Evaluates instruction following for enterprise/API pipelines โ exact output format, strict step ordering, ranking, and calibrated refusal across six categories.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
Evaluates instruction following for enterprise/API pipelines โ exact output format, strict step ordering, ranking, and calibrated refusal across six categories.
Bilingual benchmark of how models follow instructions when several constraints must hold at once โ where stacking constraints is not linear.
707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints โ whether instruction following survives deployment-scale prompts.
80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1โ10 scale.
155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).
~500 prompts, each carrying an explicit, machine-checkable constraint (format, content, or style) that a response must respect.