Logical Reasoning & Test-Time Compute
Reasoning models utilize reinforcement learning with verified reward signals and dynamic inference-time compute. By generating internal verification chains before answering, models achieve major leaps on previously unsolvable logical puzzles.
Forensic Analysis & Methodology
Why Classic Benchmarks Failed
Static multiple-choice benchmarks like MMLU and ARC suffered from severe web contamination, allowing models to memorize surface patterns rather than executing genuine multi-hop reasoning.
The TrustTheBench & LiveBench Approach
LiveBench Reasoning and GPQA Diamond introduce monthly refreshed questions and PhD-level expert problems with deterministic auto-grading that forbid memorization.
Verified Category Leaderboard
Ranked strictly by verified Logical Reasoning & Test-Time Compute performance.
| Rank | Model | Research Lab | Access Type | Category Score | Speed | Pricing / 1M |
|---|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | proprietary | 94.8% | 85 tok/s | $2.50 in |
| #2 | Grok 4 | xAI | proprietary | 94.5% | 75 tok/s | $5.00 in |
| #3 | Claude Sonnet 5 | Anthropic | proprietary | 93.0% | 95 tok/s | $3.00 in |
| #4 | Gemini 3 Pro | Google DeepMind | proprietary | 92.0% | 80 tok/s | $2.50 in |
| #5 | GPT-6 Astra (max) | OpenAI | proprietary | 91.7% | 62 tok/s | $4.50 in |
| #6 | Claude Opus 5 | Anthropic | proprietary | 91.2% | 48 tok/s | $5.00 in |
| #7 | Kimi K3 | Moonshot AI | open-weight | 90.7% | 58 tok/s | $2.80 in |
| #8 | Claude Fable 5 | Anthropic | proprietary | 89.7% | 45 tok/s | $10.00 in |
| #9 | Claude Fable 5.1 (Hybrid Reasoning) | Anthropic | proprietary | 89.7% | 54 tok/s | $5.00 in |
| #10 | GPT-5.5 Thinking | OpenAI | proprietary | 89.7% | 70 tok/s | $3.00 in |
| #11 | o3 | OpenAI | proprietary | 89.5% | 68 tok/s | $4.00 in |
| #12 | DeepSeek V4.1 Flash | DeepSeek | proprietary | 88.4% | 96 tok/s | $0.18 in |
| #13 | Grok 4.6 | xAI | proprietary | 88.4% | 75 tok/s | $2.00 in |
| #14 | Llama 4 Maverick | Meta AI | open-weight | 88.2% | 82 tok/s | $0.50 in |
| #15 | GPT-5.6 Sol | OpenAI | proprietary | 88.0% | 62 tok/s | $5.00 in |
Deployment Recommendations
GPT-5.6 Sol (91.7%) & Claude Fable 5 (89.7%)
Maximum reasoning depth and lowest error rate on difficult boundary problems.
Kimi K3 (90.7%) & DeepSeek-R1 (87.9%)
Host locally or via cost-effective cloud providers with full data privacy.
Gemini 3.7 Flash (87.8% at $0.75/1M) & o3-mini
Optimal price-to-performance ratio for high-volume automated agent pipelines.