IFEval
~500 prompts, each carrying an explicit, machine-checkable constraint (format, content, or style) that a response must respect.
Deprecated in this wiki for frontier instruction-following comparison: it remains a useful cheap regression check, but its simple verifiable prompts have become a floor and newer suites cover multi-constraint, multi-turn, and agentic instruction following more directly.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| DeepSeek V4-Flash | 62% | 2026-08-20 | independent | Source ↗ |
| DeepSeek V4-Pro | 64.8% | 2026-08-20 | independent | Source ↗ |
| Claude Fable 5.1 (Hybrid Reasoning) | 81.3% | 2026-08-19 | independent | Source ↗ |
| Gemini 3.7 Flash | 65.9% | 2026-08-17 | independent | Source ↗ |
| Grok 4.6 | 65.1% | 2026-08-16 | independent | Source ↗ |
| GPT-5 | 95.5% | 2026-08-11 | independent | Source ↗ |
| DeepSeek V4.1 Flash | 82% | 2026-08-05 | independent | Source ↗ |
| GPT-6 Astra (max) | 78% | 2026-08-01 | independent | Source ↗ |
| Kimi K3 | 76% | 2026-07-30 | independent | Source ↗ |
| Claude Sonnet 5 | 94.8% | 2026-07-28 | independent | Source ↗ |
| Claude Opus 5 | 67.5% | 2026-07-28 | independent | Source ↗ |
| Gemini 3.6 Flash | 64.5% | 2026-07-25 | independent | Source ↗ |
| Gemini 3.5 Flash-Lite | 61.5% | 2026-07-25 | independent | Source ↗ |
| GPT-5.5 Thinking | 68% | 2026-07-13 | independent | Source ↗ |
| GPT-5.6 Sol | 80% | 2026-07-13 | independent | Source ↗ |
| Qwen3-235B | 64% | 2026-06-22 | independent | Source ↗ |
| QwQ Plus | 63% | 2026-06-22 | independent | Source ↗ |
| Grok 4 | 94% | 2026-06-19 | independent | Source ↗ |
| GLM-5.2 | 66% | 2026-06-16 | independent | Source ↗ |
| Claude Fable 5 | 70% | 2026-06-13 | independent | Source ↗ |
| Gemini 3 Pro | 93.5% | 2026-05-22 | independent | Source ↗ |
| o3 | 77.2% | 2025-12-24 | independent | Source ↗ |
| Mistral Large 3 | 88% | 2025-11-16 | independent | Source ↗ |
| Gemma 3 27B Preview | 74% | 2025-04-16 | independent | Source ↗ |
| Llama 4 Maverick | 90% | 2025-04-09 | independent | Source ↗ |
| Llama 4 Scout | 89% | 2025-04-09 | independent | Source ↗ |
| Gemini 2.5 Pro | 77% | 2025-03-29 | independent | Source ↗ |
| QwQ 32B | 85% | 2025-03-09 | independent | Source ↗ |
| GPT-4.5 | 91.5% | 2025-03-03 | independent | Source ↗ |
| Claude 3.7 Sonnet | 77% | 2025-02-28 | independent | Source ↗ |
| Grok 3 | 76.5% | 2025-02-21 | independent | Source ↗ |
| Gemini 2.5 Flash (Thinking) | 77% | 2025-02-14 | independent | Source ↗ |
| Gemini 2.0 Pro | 88% | 2025-02-09 | independent | Source ↗ |
| Sonar Reasoning Pro | 75.2% | 2025-02-09 | independent | Source ↗ |
| Gemini 2.0 Flash | 88.7% | 2025-02-09 | independent | Source ↗ |
| o3-mini | 75% | 2025-02-04 | independent | Source ↗ |
| Qwen 2.5 Max | 87% | 2025-02-01 | independent | Source ↗ |
| DeepSeek-R1 | 75.2% | 2025-01-24 | independent | Source ↗ |
| Gemini 2.0 Flash Thinking | 89% | 2025-01-24 | independent | Source ↗ |
| DeepSeek-V3 | 87.5% | 2024-12-30 | independent | Source ↗ |
| Sonar Pro | 75.5% | 2024-12-14 | independent | Source ↗ |
| Llama 3.3 70B | 86.2% | 2024-12-10 | independent | Source ↗ |
| Llama 3.3 70B Instruct | 72.5% | 2024-12-10 | independent | Source ↗ |
| o1 | 87.5% | 2024-12-09 | independent | Source ↗ |
| QwQ-32B Preview | 75% | 2024-12-02 | independent | Source ↗ |
| GPT-4o | 90.8% | 2024-11-24 | independent | Source ↗ |
| Sonar | 73% | 2024-11-19 | independent | Source ↗ |
| Claude 3.5 Sonnet (v2) | 75% | 2024-10-26 | independent | Source ↗ |
| Claude 3.5 Sonnet | 89.5% | 2024-10-26 | independent | Source ↗ |
| Qwen 2.5 72B Instruct | 72% | 2024-09-23 | independent | Source ↗ |
| Gemma 2 2B | 60% | 2024-08-04 | independent | Source ↗ |
| Mistral Large 2 | 72% | 2024-07-28 | independent | Source ↗ |
| Llama 3.1 405B | 86.5% | 2024-07-27 | independent | Source ↗ |
| Gemma 2 27B | 71% | 2024-07-01 | independent | Source ↗ |
| Gemma 2 9B | 67.5% | 2024-07-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Measured human strict compliance rate under complex multi-rule constraints (word counts, punctuation, paragraph limits).
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.prompt- and instruction-level accuracy, strict and loose (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks ifeval --batch_size autoopencompass --datasets ifeval --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for IFEval, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace