Complex Rule Compliance, Word Bounds & Formatting Enforcement

Instruction Following & Negative Constraints

Testing whether models obey strict negative constraints ('Do not include letters X', 'Output exactly 3 bullet points in lowercase JSON') across multi-turn prompts.

Human Baseline92.5% (Precise Compliance)
Contamination StatusLow (Deterministic constraint checking)
Top Frontier ModelClaude Fable 5 (70.0%) & GPT-5.6 Sol (68.5%)

Forensic Analysis & Methodology

Why Classic Benchmarks Failed

Subjective LLM-as-a-judge evaluations suffered from length bias, favoring verbose paragraphs even when the user explicitly asked for brevity.

The TrustTheBench & LiveBench Approach

IFEval and LiveBench IF execute programmatic assertions (regex, string counts, JSON validators) to measure deterministic compliance.

Verified Category Leaderboard

Ranked strictly by verified Instruction Following & Negative Constraints performance.

RankModelResearch LabAccess TypeCategory ScoreSpeedPricing / 1M
#1GPT-5OpenAIproprietary95.5%85 tok/s$2.50 in
#2Claude Sonnet 5Anthropicproprietary94.8%95 tok/s$3.00 in
#3Grok 4xAIproprietary94.0%75 tok/s$5.00 in
#4Gemini 3 ProGoogle DeepMindproprietary93.5%80 tok/s$2.50 in
#5GPT-4.5OpenAIproprietary91.5%42 tok/s$75.00 in
#6GPT-4oOpenAIproprietary90.8%70 tok/s$2.50 in
#7Llama 4 MaverickMeta AIopen-weight90.0%82 tok/s$0.50 in
#8Claude 3.5 SonnetAnthropicproprietary89.5%75 tok/s$3.00 in
#9Llama 4 ScoutMeta AIopen-weight89.0%140 tok/s$0.20 in
#10Gemini 2.0 Flash ThinkingGoogle DeepMindproprietary89.0%95 tok/s$0.10 in
#11Gemini 2.0 FlashGoogle DeepMindproprietary88.7%185 tok/s$0.10 in
#12Gemini 2.0 ProGoogle DeepMindproprietary88.0%60 tok/s$1.50 in
#13Mistral Large 3Mistral AIproprietary88.0%60 tok/s$2.00 in
#14o1OpenAIproprietary87.5%45 tok/s$15.00 in
#15DeepSeek-V3DeepSeekopen-weight87.5%62 tok/s$0.14 in

Deployment Recommendations

🏆 Best Overall Frontier

Claude Fable 5 (70.0%) & GPT-5.6 Sol (68.5%)

Maximum reasoning depth and lowest error rate on difficult boundary problems.

🔓 Best Open-Weight

Llama 4 Maverick (90.0%) & Llama 4 Scout (89.0%)

Host locally or via cost-effective cloud providers with full data privacy.

⚡ Best Value & Speed

Claude 3.7 Sonnet ($3.00/1M)

Optimal price-to-performance ratio for high-volume automated agent pipelines.

Relevant Benchmark References

IFEval

Metric: Strict Accuracy %View Methodology →

LiveBench IF

Metric: Constraint Pass %View Methodology →