Tabular Calculations, JSON Schema Validation & Multi-Hop Retrieval
Data Analysis & Structured Reasoning
Evaluating how accurately foundation models interpret unstructured tables, compute financial or statistical metrics, and format strict JSON schemas without hallucinated keys.
Human Baseline75.0% (Data Analyst Baseline)
Contamination StatusModerate on static tables; Low on dynamic synthesis
Top Frontier ModelClaude Fable 5 (81.2%) & Gemini 3 Pro (88.0%)
Forensic Analysis & Methodology
Why Classic Benchmarks Failed
Standard Q&A benchmarks tested text extraction rather than actual numerical aggregation, multi-table joins, and schema enforcement.
The TrustTheBench & LiveBench Approach
LiveBench Data Analysis and TableMWP require models to parse complex CSV/JSON structures and execute programmatic data transformations.
Verified Category Leaderboard
Ranked strictly by verified Data Analysis & Structured Reasoning performance.
| Rank | Model | Research Lab | Access Type | Category Score | Speed | Pricing / 1M |
|---|---|---|---|---|---|---|
| #1 | GPT-5 | OpenAI | proprietary | 88.5% | 85 tok/s | $2.50 in |
| #2 | Gemini 3 Pro | Google DeepMind | proprietary | 88.0% | 80 tok/s | $2.50 in |
| #3 | Grok 4 | xAI | proprietary | 87.0% | 75 tok/s | $5.00 in |
| #4 | Claude Sonnet 5 | Anthropic | proprietary | 86.5% | 95 tok/s | $3.00 in |
| #5 | Claude Fable 5.1 (Hybrid Reasoning) | Anthropic | proprietary | 82.1% | 54 tok/s | $5.00 in |
| #6 | Llama 4 Maverick | Meta AI | open-weight | 82.0% | 82 tok/s | $0.50 in |
| #7 | GPT-4.5 | OpenAI | proprietary | 81.5% | 42 tok/s | $75.00 in |
| #8 | GPT-6 Astra (max) | OpenAI | proprietary | 81.4% | 62 tok/s | $4.50 in |
| #9 | Claude Fable 5 | Anthropic | proprietary | 81.2% | 45 tok/s | $10.00 in |
| #10 | DeepSeek V4.1 Flash | DeepSeek | proprietary | 80.5% | 96 tok/s | $0.18 in |
| #11 | GPT-5.6 Sol | OpenAI | proprietary | 80.1% | 62 tok/s | $5.00 in |
| #12 | Gemini 2.0 Pro | Google DeepMind | proprietary | 80.1% | 60 tok/s | $1.50 in |
| #13 | Llama 4 Scout | Meta AI | open-weight | 80.0% | 140 tok/s | $0.20 in |
| #14 | o3 | OpenAI | proprietary | 79.5% | 68 tok/s | $4.00 in |
| #15 | Claude 3.7 Sonnet | Anthropic | proprietary | 79.2% | 72 tok/s | $3.00 in |
Deployment Recommendations
🏆 Best Overall Frontier
Claude Fable 5 (81.2%) & Gemini 3 Pro (88.0%)
Maximum reasoning depth and lowest error rate on difficult boundary problems.
🔓 Best Open-Weight
Llama 4 Maverick (82.0%) & Kimi K3 (77.2%)
Host locally or via cost-effective cloud providers with full data privacy.
⚡ Best Value & Speed
Gemini 3.5 Flash-Lite ($0.10/1M)
Optimal price-to-performance ratio for high-volume automated agent pipelines.