Tokens/Second, TTFT Latency & API Cost Efficiency
Inference Speed, Latency & Price Frontier
Benchmarking real-world generation throughput (tokens/sec), time-to-first-token (TTFT ms), and pricing per million tokens across provider APIs and open-weight hostings.
Human BaselineN/A (Hardware & Quantization Metric)
Contamination StatusZero (Empirical runtime telemetry)
Top Frontier ModelGemini 3.7 Flash (182 tok/s at $0.75/1M)
Forensic Analysis & Methodology
Why Classic Benchmarks Failed
Static model scoreboards ignored deployment latency, making models with high latency impractical for real-time agentic tool loops.
The TrustTheBench & LiveBench Approach
Artificial Analysis continuous API probing across 20+ cloud providers under standardized load conditions.
Verified Category Leaderboard
Ranked strictly by verified Inference Speed, Latency & Price Frontier performance.
| Rank | Model | Research Lab | Access Type | Category Score | Speed | Pricing / 1M |
|---|---|---|---|---|---|---|
| #1 | Gemini 3.5 Flash-Lite | Google DeepMind | proprietary | 220 tok/s | 220 tok/s | $0.10 in |
| #2 | Amazon Nova Lite | Amazon AWS | proprietary | 200 tok/s | 200 tok/s | $0.06 in |
| #3 | DeepSeek V4-Flash | DeepSeek | open-weight | 190 tok/s | 190 tok/s | $0.14 in |
| #4 | Gemini 2.0 Flash | Google DeepMind | proprietary | 185 tok/s | 185 tok/s | $0.10 in |
| #5 | Gemini 3.7 Flash | Google DeepMind | proprietary | 182 tok/s | 182 tok/s | $0.75 in |
| #6 | Gemma 2 2B | Google DeepMind | open-weight | 180 tok/s | 180 tok/s | $0.05 in |
| #7 | Gemini 3.6 Flash | Google DeepMind | proprietary | 175 tok/s | 175 tok/s | $0.75 in |
| #8 | Gemini 2.5 Flash (Thinking) | Google DeepMind | proprietary | 145 tok/s | 145 tok/s | $0.07 in |
| #9 | Llama 4 Scout | Meta AI | open-weight | 140 tok/s | 140 tok/s | $0.20 in |
| #10 | Gemini 1.5 Flash | Google DeepMind | proprietary | 140 tok/s | 140 tok/s | $0.07 in |
| #11 | Gemma 2 9B | Google DeepMind | open-weight | 130 tok/s | 130 tok/s | $0.10 in |
| #12 | GPT-4o mini | OpenAI | proprietary | 125 tok/s | 125 tok/s | $0.15 in |
| #13 | Amazon Nova Pro | Amazon AWS | proprietary | 120 tok/s | 120 tok/s | $0.80 in |
| #14 | o3-mini | OpenAI | proprietary | 115 tok/s | 115 tok/s | $1.10 in |
| #15 | Yi-Lightning | 01.AI | proprietary | 110 tok/s | 110 tok/s | $0.14 in |
Deployment Recommendations
🏆 Best Overall Frontier
Gemini 3.7 Flash (182 tok/s at $0.75/1M)
Maximum reasoning depth and lowest error rate on difficult boundary problems.
🔓 Best Open-Weight
DeepSeek V4-Flash (190 tok/s at $0.14/1M)
Host locally or via cost-effective cloud providers with full data privacy.
⚡ Best Value & Speed
Gemini 3.5 Flash-Lite (220 tok/s at $0.10/1M)
Optimal price-to-performance ratio for high-volume automated agent pipelines.