Complex Reasoning, Multi-Step Logic & Deduction

Logical Reasoning & Test-Time Compute

Reasoning models utilize reinforcement learning with verified reward signals and dynamic inference-time compute. By generating internal verification chains before answering, models achieve major leaps on previously unsolvable logical puzzles.

Human Baseline69.7% (Expert Generalist)
Contamination StatusHigh on static suites (MMLU / ARC Easy); Low on monthly LiveBench
Top Frontier ModelGPT-5.6 Sol (91.7%) & Claude Fable 5 (89.7%)

Forensic Analysis & Methodology

Why Classic Benchmarks Failed

Static multiple-choice benchmarks like MMLU and ARC suffered from severe web contamination, allowing models to memorize surface patterns rather than executing genuine multi-hop reasoning.

The TrustTheBench & LiveBench Approach

LiveBench Reasoning and GPQA Diamond introduce monthly refreshed questions and PhD-level expert problems with deterministic auto-grading that forbid memorization.

Verified Category Leaderboard

Ranked strictly by verified Logical Reasoning & Test-Time Compute performance.

RankModelResearch LabAccess TypeCategory ScoreSpeedPricing / 1M
#1GPT-5OpenAIproprietary94.8%85 tok/s$2.50 in
#2Grok 4xAIproprietary94.5%75 tok/s$5.00 in
#3Claude Sonnet 5Anthropicproprietary93.0%95 tok/s$3.00 in
#4Gemini 3 ProGoogle DeepMindproprietary92.0%80 tok/s$2.50 in
#5GPT-6 Astra (max)OpenAIproprietary91.7%62 tok/s$4.50 in
#6Claude Opus 5Anthropicproprietary91.2%48 tok/s$5.00 in
#7Kimi K3Moonshot AIopen-weight90.7%58 tok/s$2.80 in
#8Claude Fable 5Anthropicproprietary89.7%45 tok/s$10.00 in
#9Claude Fable 5.1 (Hybrid Reasoning)Anthropicproprietary89.7%54 tok/s$5.00 in
#10GPT-5.5 ThinkingOpenAIproprietary89.7%70 tok/s$3.00 in
#11o3OpenAIproprietary89.5%68 tok/s$4.00 in
#12DeepSeek V4.1 FlashDeepSeekproprietary88.4%96 tok/s$0.18 in
#13Grok 4.6xAIproprietary88.4%75 tok/s$2.00 in
#14Llama 4 MaverickMeta AIopen-weight88.2%82 tok/s$0.50 in
#15GPT-5.6 SolOpenAIproprietary88.0%62 tok/s$5.00 in

Deployment Recommendations

🏆 Best Overall Frontier

GPT-5.6 Sol (91.7%) & Claude Fable 5 (89.7%)

Maximum reasoning depth and lowest error rate on difficult boundary problems.

🔓 Best Open-Weight

Kimi K3 (90.7%) & DeepSeek-R1 (87.9%)

Host locally or via cost-effective cloud providers with full data privacy.

⚡ Best Value & Speed

Gemini 3.7 Flash (87.8% at $0.75/1M) & o3-mini

Optimal price-to-performance ratio for high-volume automated agent pipelines.

Relevant Benchmark References

LiveBench Reasoning

Metric: Accuracy %View Methodology →

GPQA Diamond

Metric: Zero-Shot Accuracy %View Methodology →

ARC-AGI-2

Metric: Pass@1 %View Methodology →

HLE (Humanity's Last Exam)

Metric: Accuracy %View Methodology →