RESEARCH & CAPABILITY OBSERVATORY

AI Evaluation & Benchmark Insights

All domain deep-dives, contamination audits, and frontier capability analyses consolidated on one page as of September 2026.

The Reasoning Paradigm Shift

Test-time compute scaling (Claude Fable 5, GPT-5.6 Sol, Kimi K3, DeepSeek-R1) decouples reasoning accuracy from static parameters, yielding record AIME and GPQA results.

The Open-Weight Value Frontier

Open-weight architectures like DeepSeek V4-Flash ($0.14/1M) and Kimi K3 deliver frontier intelligence at unprecedented cost efficiency.

Contamination-Resistant Testing

Static benchmarks (MMLU, GSM8K) have saturated from training data leaks. Monthly refreshed suites (LiveBench, SWE-bench Verified) provide the true signal.

Multimodal Capability Frontier Spider

Comparing multi-axis performance envelopes across reasoning, coding, olympiad math, and instruction fidelity.

Cost vs. Intelligence Pareto Frontier (September 2026)

Empirical quality versus API pricing ($/1M blended tokens). Models closer to the top-left provide the highest value per dollar.

Proprietary Frontier
Open-Weight Workhorse

Logical Reasoning & Test-Time Compute

Complex Reasoning, Multi-Step Logic & Deduction

Human Baseline69.7% (Expert Generalist)
Contamination RiskLow on monthly LiveBench; High on static MMLU

Why Legacy Benchmarks Saturated

Static multiple-choice benchmarks like MMLU and ARC suffered from severe web contamination, allowing models to memorize surface patterns rather than executing genuine multi-hop reasoning.

Modern Evaluation Protocol

LiveBench Reasoning and GPQA Diamond introduce monthly refreshed questions and PhD-level expert problems with deterministic auto-grading that forbid memorization.

Verified Model Score Distribution (%)Human Reference: 69.7%

Domain Leaders & Recommendations

Frontier SOTA
GPT-5.6 Sol (91.7%) & Claude Fable 5 (89.7%)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
Kimi K3 (90.7%) & DeepSeek-R1 (87.9%)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
Gemini 3.7 Flash (87.8% at $0.75/1M) & o3-mini

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Logical Reasoning Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1GPT-5OpenAI94.8%$2.50 / $10.0085 tok/s
#2Grok 4xAI94.5%$5.00 / $20.0075 tok/s
#3Claude Sonnet 5Anthropic93.0%$3.00 / $15.0095 tok/s
#4Gemini 3 ProGoogle DeepMind92.0%$2.50 / $10.0080 tok/s
#5GPT-6 Astra (max)OpenAI91.7%$4.50 / $20.0062 tok/s
#6Claude Opus 5Anthropic91.2%$5.00 / $25.0048 tok/s
#7Kimi K3Moonshot AI90.7%$2.80 / $14.0058 tok/s
#8Claude Fable 5Anthropic89.7%$10.00 / $50.0045 tok/s

Software Engineering & Agentic Coding

Repository-Level Debugging, Code Synthesis & Refactoring

Human Baseline78.4% (Senior SWE Baseline)
Contamination RiskControlled on SWE-bench Verified; Critical on HumanEval

Why Legacy Benchmarks Saturated

HumanEval and MBPP consist of trivial single-function Python prompts that exist in hundreds of thousands of public training files, saturating over 90% and failing to predict real-world developer utility.

Modern Evaluation Protocol

SWE-bench Verified and LiveCodeBench evaluate real pull requests and newly released LeetCode/Codeforces problems with hidden test suites.

Verified Model Score Distribution (%)Human Reference: 78.4%

Domain Leaders & Recommendations

Frontier SOTA
Claude Fable 5 (86.0%) & GPT-5.6 Sol (83.9%)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
Kimi K3 (81.4%) & DeepSeek V4-Pro (81.0%)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
DeepSeek V4-Flash ($0.14/1M) & Qwen 2.5 Coder 32B

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Software Engineering Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1Claude Fable 5.1 (Hybrid Reasoning)Anthropic86.4%$5.00 / $25.0054 tok/s
#2Claude Fable 5Anthropic86.0%$10.00 / $50.0045 tok/s
#3GPT-6 Astra (max)OpenAI85.9%$4.50 / $20.0062 tok/s
#4DeepSeek V4.1 FlashDeepSeek85.1%$0.18 / $0.5596 tok/s
#5GPT-5OpenAI84.5%$2.50 / $10.0085 tok/s
#6GPT-5.6 SolOpenAI84.1%$5.00 / $30.0062 tok/s
#7Grok 4xAI84.0%$5.00 / $20.0075 tok/s
#8o3OpenAI84.0%$4.00 / $18.0068 tok/s

Mathematics & Olympiad Problem Solving

AIME, Olympiad Proofs & Formal Mathematics

Human Baseline40.0% (USAMO Qualifier Baseline)
Contamination RiskLow on FrontierMath & AIME 2026; Total on GSM8K

Why Legacy Benchmarks Saturated

GSM8K reached 97%+ saturation across virtually all frontier models due to verbatim training data inclusion and low problem complexity.

Modern Evaluation Protocol

FrontierMath and newly released AIME 2025/2026 exams test rigorous multi-step numerical calculation and symbolic deduction with near-zero pre-training contamination.

Verified Model Score Distribution (%)Human Reference: 40%

Domain Leaders & Recommendations

Frontier SOTA
GPT-5.6 Sol (96.2%) & Claude Fable 5 (96.0%)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
QwQ Plus (92.5%) & DeepSeek-R1 (90.8%)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
Gemini 3.7 Flash (93.5% at $0.75/1M)

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Mathematics Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1Claude Fable 5Anthropic96.0%$10.00 / $50.0045 tok/s
#2GPT-5.5 ThinkingOpenAI95.9%$3.00 / $15.0070 tok/s
#3Claude Opus 5Anthropic95.7%$5.00 / $25.0048 tok/s
#4Grok 4xAI95.0%$5.00 / $20.0075 tok/s
#5GPT-5OpenAI94.0%$2.50 / $10.0085 tok/s
#6Gemini 3.7 FlashGoogle DeepMind93.5%$0.75 / $3.75182 tok/s
#7QwQ PlusAlibaba Cloud / Qwen92.5%$0.80 / $2.4052 tok/s
#8Grok 4.6xAI92.1%$2.00 / $6.0075 tok/s
Audited Benchmark Suites:

Data Analysis & Structured Reasoning

Tabular Calculations, JSON Schema Validation & Multi-Hop Retrieval

Human Baseline75.0% (Data Analyst Baseline)
Contamination RiskModerate on static tables; Low on dynamic synthesis

Why Legacy Benchmarks Saturated

Standard Q&A benchmarks tested text extraction rather than actual numerical aggregation, multi-table joins, and schema enforcement.

Modern Evaluation Protocol

LiveBench Data Analysis and TableMWP require models to parse complex CSV/JSON structures and execute programmatic data transformations.

Verified Model Score Distribution (%)Human Reference: 75%

Domain Leaders & Recommendations

Frontier SOTA
Claude Fable 5 (81.2%) & Gemini 3 Pro (88.0%)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
Llama 4 Maverick (82.0%) & Kimi K3 (77.2%)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
Gemini 3.5 Flash-Lite ($0.10/1M)

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Data Analysis Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1GPT-5OpenAI88.5%$2.50 / $10.0085 tok/s
#2Gemini 3 ProGoogle DeepMind88.0%$2.50 / $10.0080 tok/s
#3Grok 4xAI87.0%$5.00 / $20.0075 tok/s
#4Claude Sonnet 5Anthropic86.5%$3.00 / $15.0095 tok/s
#5Claude Fable 5.1 (Hybrid Reasoning)Anthropic82.1%$5.00 / $25.0054 tok/s
#6Llama 4 MaverickMeta AI82.0%$0.50 / $1.5082 tok/s
#7GPT-4.5OpenAI81.5%$75.00 / $150.0042 tok/s
#8GPT-6 Astra (max)OpenAI81.4%$4.50 / $20.0062 tok/s
Audited Benchmark Suites:

Instruction Following & Negative Constraints

Complex Rule Compliance, Word Bounds & Formatting Enforcement

Human Baseline92.5% (Precise Compliance)
Contamination RiskLow (Deterministic constraint checking)

Why Legacy Benchmarks Saturated

Subjective LLM-as-a-judge evaluations suffered from length bias, favoring verbose paragraphs even when the user explicitly asked for brevity.

Modern Evaluation Protocol

IFEval and LiveBench IF execute programmatic assertions (regex, string counts, JSON validators) to measure deterministic compliance.

Verified Model Score Distribution (%)Human Reference: 92.5%

Domain Leaders & Recommendations

Frontier SOTA
Claude Fable 5 (70.0%) & GPT-5.6 Sol (68.5%)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
Llama 4 Maverick (90.0%) & Llama 4 Scout (89.0%)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
Claude 3.7 Sonnet ($3.00/1M)

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Instruction Following Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1GPT-5OpenAI95.5%$2.50 / $10.0085 tok/s
#2Claude Sonnet 5Anthropic94.8%$3.00 / $15.0095 tok/s
#3Grok 4xAI94.0%$5.00 / $20.0075 tok/s
#4Gemini 3 ProGoogle DeepMind93.5%$2.50 / $10.0080 tok/s
#5GPT-4.5OpenAI91.5%$75.00 / $150.0042 tok/s
#6GPT-4oOpenAI90.8%$2.50 / $10.0070 tok/s
#7Llama 4 MaverickMeta AI90.0%$0.50 / $1.5082 tok/s
#8Claude 3.5 SonnetAnthropic89.5%$3.00 / $15.0075 tok/s
Audited Benchmark Suites:

Inference Speed, Latency & Price Frontier

Tokens/Sec Throughput, TTFT Latency & API Cost Efficiency

Human BaselineN/A (Hardware Bound)
Contamination RiskZero (Direct API telemetry)

Why Legacy Benchmarks Saturated

Classic evaluation ignored inference cost and latency, ranking massive multi-trillion parameter clusters alongside lightweight edge models without normalization.

Modern Evaluation Protocol

Artificial Analysis speed testing and standardized blended pricing metrics reveal true cost-per-task efficiency.

Inference Speed (tok/s) vs Output Price ($/1M tokens)Pareto Frontier Curve

Domain Leaders & Recommendations

Frontier SOTA
Gemini 3.7 Flash (182 tok/s at $0.75/1M)

Top absolute accuracy and benchmark score ceiling.

Open-Weight SOTA
DeepSeek V4-Flash (190 tok/s at $0.14/1M)

Best self-hostable open model weights and weights-available checkpoints.

Value / Speed Champion
Amazon Nova Lite (200 tok/s at $0.06/1M)

Optimal quality-to-cost ratio for high-volume automated production pipelines.

Verified Inference Speed, Latency Leaderboard

RankModelCompanyVerified ScorePricing (In/Out)Speed
#1Gemini 3.5 Flash-LiteGoogle DeepMind220 tok/s$0.10 / $0.40220 tok/s
#2Amazon Nova LiteAmazon AWS200 tok/s$0.06 / $0.24200 tok/s
#3DeepSeek V4-FlashDeepSeek190 tok/s$0.14 / $0.28190 tok/s
#4Gemini 2.0 FlashGoogle DeepMind185 tok/s$0.10 / $0.40185 tok/s
#5Gemini 3.7 FlashGoogle DeepMind182 tok/s$0.75 / $3.75182 tok/s
#6Gemma 2 2BGoogle DeepMind180 tok/s$0.05 / $0.10180 tok/s
#7Gemini 3.6 FlashGoogle DeepMind175 tok/s$0.75 / $3.75175 tok/s
#8Gemini 2.5 Flash (Thinking)Google DeepMind145 tok/s$0.07 / $0.30145 tok/s