AI Evaluation & Benchmark Insights
All domain deep-dives, contamination audits, and frontier capability analyses consolidated on one page as of September 2026.
The Reasoning Paradigm Shift
Test-time compute scaling (Claude Fable 5, GPT-5.6 Sol, Kimi K3, DeepSeek-R1) decouples reasoning accuracy from static parameters, yielding record AIME and GPQA results.
The Open-Weight Value Frontier
Open-weight architectures like DeepSeek V4-Flash ($0.14/1M) and Kimi K3 deliver frontier intelligence at unprecedented cost efficiency.
Contamination-Resistant Testing
Static benchmarks (MMLU, GSM8K) have saturated from training data leaks. Monthly refreshed suites (LiveBench, SWE-bench Verified) provide the true signal.
Multimodal Capability Frontier Spider
Comparing multi-axis performance envelopes across reasoning, coding, olympiad math, and instruction fidelity.
Cost vs. Intelligence Pareto Frontier (September 2026)
Empirical quality versus API pricing ($/1M blended tokens). Models closer to the top-left provide the highest value per dollar.
Logical Reasoning & Test-Time Compute
Complex Reasoning, Multi-Step Logic & Deduction
Why Legacy Benchmarks Saturated
Static multiple-choice benchmarks like MMLU and ARC suffered from severe web contamination, allowing models to memorize surface patterns rather than executing genuine multi-hop reasoning.
Modern Evaluation Protocol
LiveBench Reasoning and GPQA Diamond introduce monthly refreshed questions and PhD-level expert problems with deterministic auto-grading that forbid memorization.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Logical Reasoning Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | 94.8% | $2.50 / $10.00 | 85 tok/s |
| #2 | Grok 4 | xAI | 94.5% | $5.00 / $20.00 | 75 tok/s |
| #3 | Claude Sonnet 5 | Anthropic | 93.0% | $3.00 / $15.00 | 95 tok/s |
| #4 | Gemini 3 Pro | Google DeepMind | 92.0% | $2.50 / $10.00 | 80 tok/s |
| #5 | GPT-6 Astra (max) | OpenAI | 91.7% | $4.50 / $20.00 | 62 tok/s |
| #6 | Claude Opus 5 | Anthropic | 91.2% | $5.00 / $25.00 | 48 tok/s |
| #7 | Kimi K3 | Moonshot AI | 90.7% | $2.80 / $14.00 | 58 tok/s |
| #8 | Claude Fable 5 | Anthropic | 89.7% | $10.00 / $50.00 | 45 tok/s |
Software Engineering & Agentic Coding
Repository-Level Debugging, Code Synthesis & Refactoring
Why Legacy Benchmarks Saturated
HumanEval and MBPP consist of trivial single-function Python prompts that exist in hundreds of thousands of public training files, saturating over 90% and failing to predict real-world developer utility.
Modern Evaluation Protocol
SWE-bench Verified and LiveCodeBench evaluate real pull requests and newly released LeetCode/Codeforces problems with hidden test suites.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Software Engineering Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | Claude Fable 5.1 (Hybrid Reasoning) | Anthropic | 86.4% | $5.00 / $25.00 | 54 tok/s |
| #2 | Claude Fable 5 | Anthropic | 86.0% | $10.00 / $50.00 | 45 tok/s |
| #3 | GPT-6 Astra (max) | OpenAI | 85.9% | $4.50 / $20.00 | 62 tok/s |
| #4 | DeepSeek V4.1 Flash | DeepSeek | 85.1% | $0.18 / $0.55 | 96 tok/s |
| #5 | GPT-5 | OpenAI | 84.5% | $2.50 / $10.00 | 85 tok/s |
| #6 | GPT-5.6 Sol | OpenAI | 84.1% | $5.00 / $30.00 | 62 tok/s |
| #7 | Grok 4 | xAI | 84.0% | $5.00 / $20.00 | 75 tok/s |
| #8 | o3 | OpenAI | 84.0% | $4.00 / $18.00 | 68 tok/s |
Mathematics & Olympiad Problem Solving
AIME, Olympiad Proofs & Formal Mathematics
Why Legacy Benchmarks Saturated
GSM8K reached 97%+ saturation across virtually all frontier models due to verbatim training data inclusion and low problem complexity.
Modern Evaluation Protocol
FrontierMath and newly released AIME 2025/2026 exams test rigorous multi-step numerical calculation and symbolic deduction with near-zero pre-training contamination.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Mathematics Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | Claude Fable 5 | Anthropic | 96.0% | $10.00 / $50.00 | 45 tok/s |
| #2 | GPT-5.5 Thinking | OpenAI | 95.9% | $3.00 / $15.00 | 70 tok/s |
| #3 | Claude Opus 5 | Anthropic | 95.7% | $5.00 / $25.00 | 48 tok/s |
| #4 | Grok 4 | xAI | 95.0% | $5.00 / $20.00 | 75 tok/s |
| #5 | GPT-5 | OpenAI | 94.0% | $2.50 / $10.00 | 85 tok/s |
| #6 | Gemini 3.7 Flash | Google DeepMind | 93.5% | $0.75 / $3.75 | 182 tok/s |
| #7 | QwQ Plus | Alibaba Cloud / Qwen | 92.5% | $0.80 / $2.40 | 52 tok/s |
| #8 | Grok 4.6 | xAI | 92.1% | $2.00 / $6.00 | 75 tok/s |
Data Analysis & Structured Reasoning
Tabular Calculations, JSON Schema Validation & Multi-Hop Retrieval
Why Legacy Benchmarks Saturated
Standard Q&A benchmarks tested text extraction rather than actual numerical aggregation, multi-table joins, and schema enforcement.
Modern Evaluation Protocol
LiveBench Data Analysis and TableMWP require models to parse complex CSV/JSON structures and execute programmatic data transformations.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Data Analysis Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | 88.5% | $2.50 / $10.00 | 85 tok/s |
| #2 | Gemini 3 Pro | Google DeepMind | 88.0% | $2.50 / $10.00 | 80 tok/s |
| #3 | Grok 4 | xAI | 87.0% | $5.00 / $20.00 | 75 tok/s |
| #4 | Claude Sonnet 5 | Anthropic | 86.5% | $3.00 / $15.00 | 95 tok/s |
| #5 | Claude Fable 5.1 (Hybrid Reasoning) | Anthropic | 82.1% | $5.00 / $25.00 | 54 tok/s |
| #6 | Llama 4 Maverick | Meta AI | 82.0% | $0.50 / $1.50 | 82 tok/s |
| #7 | GPT-4.5 | OpenAI | 81.5% | $75.00 / $150.00 | 42 tok/s |
| #8 | GPT-6 Astra (max) | OpenAI | 81.4% | $4.50 / $20.00 | 62 tok/s |
Instruction Following & Negative Constraints
Complex Rule Compliance, Word Bounds & Formatting Enforcement
Why Legacy Benchmarks Saturated
Subjective LLM-as-a-judge evaluations suffered from length bias, favoring verbose paragraphs even when the user explicitly asked for brevity.
Modern Evaluation Protocol
IFEval and LiveBench IF execute programmatic assertions (regex, string counts, JSON validators) to measure deterministic compliance.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Instruction Following Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | 95.5% | $2.50 / $10.00 | 85 tok/s |
| #2 | Claude Sonnet 5 | Anthropic | 94.8% | $3.00 / $15.00 | 95 tok/s |
| #3 | Grok 4 | xAI | 94.0% | $5.00 / $20.00 | 75 tok/s |
| #4 | Gemini 3 Pro | Google DeepMind | 93.5% | $2.50 / $10.00 | 80 tok/s |
| #5 | GPT-4.5 | OpenAI | 91.5% | $75.00 / $150.00 | 42 tok/s |
| #6 | GPT-4o | OpenAI | 90.8% | $2.50 / $10.00 | 70 tok/s |
| #7 | Llama 4 Maverick | Meta AI | 90.0% | $0.50 / $1.50 | 82 tok/s |
| #8 | Claude 3.5 Sonnet | Anthropic | 89.5% | $3.00 / $15.00 | 75 tok/s |
Inference Speed, Latency & Price Frontier
Tokens/Sec Throughput, TTFT Latency & API Cost Efficiency
Why Legacy Benchmarks Saturated
Classic evaluation ignored inference cost and latency, ranking massive multi-trillion parameter clusters alongside lightweight edge models without normalization.
Modern Evaluation Protocol
Artificial Analysis speed testing and standardized blended pricing metrics reveal true cost-per-task efficiency.
Domain Leaders & Recommendations
Top absolute accuracy and benchmark score ceiling.
Best self-hostable open model weights and weights-available checkpoints.
Optimal quality-to-cost ratio for high-volume automated production pipelines.
Verified Inference Speed, Latency Leaderboard
| Rank | Model | Company | Verified Score | Pricing (In/Out) | Speed |
|---|---|---|---|---|---|
| #1 | Gemini 3.5 Flash-Lite | Google DeepMind | 220 tok/s | $0.10 / $0.40 | 220 tok/s |
| #2 | Amazon Nova Lite | Amazon AWS | 200 tok/s | $0.06 / $0.24 | 200 tok/s |
| #3 | DeepSeek V4-Flash | DeepSeek | 190 tok/s | $0.14 / $0.28 | 190 tok/s |
| #4 | Gemini 2.0 Flash | Google DeepMind | 185 tok/s | $0.10 / $0.40 | 185 tok/s |
| #5 | Gemini 3.7 Flash | Google DeepMind | 182 tok/s | $0.75 / $3.75 | 182 tok/s |
| #6 | Gemma 2 2B | Google DeepMind | 180 tok/s | $0.05 / $0.10 | 180 tok/s |
| #7 | Gemini 3.6 Flash | Google DeepMind | 175 tok/s | $0.75 / $3.75 | 175 tok/s |
| #8 | Gemini 2.5 Flash (Thinking) | Google DeepMind | 145 tok/s | $0.07 / $0.30 | 145 tok/s |