Complex Rule Compliance, Word Bounds & Formatting Enforcement
Instruction Following & Negative Constraints
Testing whether models obey strict negative constraints ('Do not include letters X', 'Output exactly 3 bullet points in lowercase JSON') across multi-turn prompts.
Human Baseline92.5% (Precise Compliance)
Contamination StatusLow (Deterministic constraint checking)
Top Frontier ModelClaude Fable 5 (70.0%) & GPT-5.6 Sol (68.5%)
Forensic Analysis & Methodology
Why Classic Benchmarks Failed
Subjective LLM-as-a-judge evaluations suffered from length bias, favoring verbose paragraphs even when the user explicitly asked for brevity.
The TrustTheBench & LiveBench Approach
IFEval and LiveBench IF execute programmatic assertions (regex, string counts, JSON validators) to measure deterministic compliance.
Verified Category Leaderboard
Ranked strictly by verified Instruction Following & Negative Constraints performance.
| Rank | Model | Research Lab | Access Type | Category Score | Speed | Pricing / 1M |
|---|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | proprietary | 95.5% | 85 tok/s | $2.50 in |
| #2 | Claude Sonnet 5 | Anthropic | proprietary | 94.8% | 95 tok/s | $3.00 in |
| #3 | Grok 4 | xAI | proprietary | 94.0% | 75 tok/s | $5.00 in |
| #4 | Gemini 3 Pro | Google DeepMind | proprietary | 93.5% | 80 tok/s | $2.50 in |
| #5 | GPT-4.5 | OpenAI | proprietary | 91.5% | 42 tok/s | $75.00 in |
| #6 | GPT-4o | OpenAI | proprietary | 90.8% | 70 tok/s | $2.50 in |
| #7 | Llama 4 Maverick | Meta AI | open-weight | 90.0% | 82 tok/s | $0.50 in |
| #8 | Claude 3.5 Sonnet | Anthropic | proprietary | 89.5% | 75 tok/s | $3.00 in |
| #9 | Llama 4 Scout | Meta AI | open-weight | 89.0% | 140 tok/s | $0.20 in |
| #10 | Gemini 2.0 Flash Thinking | Google DeepMind | proprietary | 89.0% | 95 tok/s | $0.10 in |
| #11 | Gemini 2.0 Flash | Google DeepMind | proprietary | 88.7% | 185 tok/s | $0.10 in |
| #12 | Gemini 2.0 Pro | Google DeepMind | proprietary | 88.0% | 60 tok/s | $1.50 in |
| #13 | Mistral Large 3 | Mistral AI | proprietary | 88.0% | 60 tok/s | $2.00 in |
| #14 | o1 | OpenAI | proprietary | 87.5% | 45 tok/s | $15.00 in |
| #15 | DeepSeek-V3 | DeepSeek | open-weight | 87.5% | 62 tok/s | $0.14 in |
Deployment Recommendations
🏆 Best Overall Frontier
Claude Fable 5 (70.0%) & GPT-5.6 Sol (68.5%)
Maximum reasoning depth and lowest error rate on difficult boundary problems.
🔓 Best Open-Weight
Llama 4 Maverick (90.0%) & Llama 4 Scout (89.0%)
Host locally or via cost-effective cloud providers with full data privacy.
⚡ Best Value & Speed
Claude 3.7 Sonnet ($3.00/1M)
Optimal price-to-performance ratio for high-volume automated agent pipelines.