AI Benchmark Leaderboard & Top Models Rankings

Verified evaluations across LiveBench contamination-free benchmarks, Artificial Analysis intelligence indices, and Autonomous Coding Agents as of September 2026.

LiveBench Contamination-Resistant Suite: Hard questions updated periodically with strict test-set isolation to prevent pre-training data memorization.

RankModel & Lab Global Avg Reasoning Coding Math Data IF Throughput Pricing / 1M
#1Claude Fable 5.1 (Hybrid Reasoning)
Anthropic★ Frontier SOTA
84.7%
89.7%
86.4%
88.5%
82.1%
81.3%
54 tok/s
$5.00/$25.00
#2GPT-6 Astra (max)
OpenAI★ Frontier SOTA
84.2%
91.7%
85.9%
90.2%
81.4%
78.0%
62 tok/s
$4.50/$20.00
#3DeepSeek V4.1 Flash
DeepSeek★ Frontier SOTA
83.9%
88.4%
85.1%
89.0%
80.5%
82.0%
96 tok/s
$0.18/$0.55
#4Claude Fable 5
Anthropic★ Frontier SOTA
83.0%
89.7%
86.0%
96.0%
81.2%
70.0%
45 tok/s
$10.00/$50.00
#5GPT-5.6 Sol
OpenAI★ Frontier SOTA
82.8%
88.0%
84.1%
87.5%
80.1%
80.0%
62 tok/s
$5.00/$30.00
#6o3
OpenAI★ Frontier SOTA
82.5%
89.5%
84.0%
88.9%
79.5%
77.2%
68 tok/s
$4.00/$18.00
#7GPT-5
OpenAI★ Frontier SOTA
82.0%
94.8%
84.5%
94.0%
88.5%
95.5%
85 tok/s
$2.50/$10.00
#8Kimi K3
Moonshot AI🔓 Open Weights
82.0%
90.7%
83.5%
87.0%
79.0%
76.0%
58 tok/s
$2.80/$14.00
#9Claude 3.7 Sonnet
Anthropic★ Frontier SOTA
81.5%
86.8%
83.2%
86.4%
79.2%
77.0%
72 tok/s
$3.00/$15.00
#10Grok 4
xAI★ Frontier SOTA
81.2%
94.5%
84.0%
95.0%
87.0%
94.0%
75 tok/s
$5.00/$20.00
#11Grok 3
xAI★ Frontier SOTA
81.2%
87.0%
82.5%
86.2%
78.0%
76.5%
65 tok/s
$3.00/$15.00
#12Gemini 2.5 Pro
Google DeepMind★ Frontier SOTA
81.0%
86.5%
82.0%
86.0%
78.5%
77.0%
85 tok/s
$1.25/$5.00
#13DeepSeek-R1
DeepSeek🔓 Open Weights
80.8%
87.9%
81.5%
87.2%
78.0%
75.2%
65 tok/s
$0.55/$2.19
#14Sonar Reasoning Pro
Perplexity AI★ Frontier SOTA
80.5%
87.0%
81.2%
86.0%
78.0%
75.2%
78 tok/s
$2.00/$8.00
#15Claude Sonnet 5
Anthropic★ Frontier SOTA
80.4%
93.0%
83.2%
91.5%
86.5%
94.8%
95 tok/s
$3.00/$15.00
#16GPT-5.5 Thinking
OpenAI★ Frontier SOTA
80.2%
89.7%
82.1%
95.9%
78.4%
68.0%
70 tok/s
$3.00/$15.00
#17o3-mini
OpenAI★ Frontier SOTA
80.2%
87.2%
81.0%
86.5%
77.0%
75.0%
115 tok/s
$1.10/$4.40
#18Claude Opus 5
Anthropic★ Frontier SOTA
80.1%
91.2%
81.4%
95.7%
78.0%
67.5%
48 tok/s
$5.00/$25.00
#19Ollama DeepSeek-R1 (Q4_K_M)
Ollama (Inference Runtime)🔓 Open Weights
80.0%
24 tok/sOpen-Weights
#20Gemini 3 Pro
Google DeepMind★ Frontier SOTA
79.8%
92.0%
81.5%
92.0%
88.0%
93.5%
80 tok/s
$2.50/$10.00
#21Gemini 2.5 Flash (Thinking)
Google DeepMindCloud API
79.8%
85.2%
80.5%
84.8%
76.5%
77.0%
145 tok/s
$0.07/$0.30
#22QwQ-32B Preview
Alibaba Cloud / QwenCloud API
79.2%
86.0%
80.0%
85.5%
76.0%
75.0%
Open-Weights
#23Gemini 3.7 Flash
Google DeepMind★ Frontier SOTA
78.8%
87.8%
78.9%
93.5%
76.5%
65.9%
182 tok/s
$0.75/$3.75
#24Grok 4.6
xAI★ Frontier SOTA
78.5%
88.4%
80.2%
92.1%
75.8%
65.1%
75 tok/s
$2.00/$6.00
#25Claude 3.5 Sonnet (v2)
AnthropicCloud API
78.5%
83.5%
79.2%
82.0%
76.0%
75.0%
Open-Weights
#26DeepSeek V4-Pro
DeepSeek🔓 Open Weights
77.9%
87.1%
81.0%
91.4%
75.2%
64.8%
65 tok/s
$0.66/$1.98
#27Qwen3-235B
Alibaba Cloud / Qwen🔓 Open Weights
77.4%
87.5%
80.5%
91.0%
74.8%
64.0%
70 tok/s
$0.80/$2.40
#28Gemini 3.6 Flash
Google DeepMind★ Frontier SOTA
77.2%
86.5%
77.8%
91.2%
75.0%
64.5%
175 tok/s
$0.75/$3.75
#29Llama 4 Maverick
Meta AI🔓 Open Weights
77.1%
88.2%
79.0%
89.5%
82.0%
90.0%
82 tok/s
$0.50/$1.50
#30GLM-5.2
Zhipu AI★ Frontier SOTA
77.0%
86.8%
79.5%
90.5%
76.0%
66.0%
72 tok/s
$1.50/$6.00
#31Sonar Pro
Perplexity AIFlagship
77.0%
81.5%
76.5%
80.0%
74.5%
75.5%
85 tok/s
$1.00/$1.00
#32MiniMax M3
MiniMax★ Frontier SOTA
76.8%
88 tok/s
$0.30/$1.20
#33QwQ Plus
Alibaba Cloud / Qwen🔓 Open Weights
76.5%
88.0%
79.0%
92.5%
74.0%
63.0%
52 tok/s
$0.80/$2.40
#34Qwen 2.5 72B Instruct
Alibaba Cloud / Qwen🔓 Open Weights
76.5%
81.0%
77.0%
81.5%
74.0%
72.0%
82 tok/s
$0.35/$0.70
#35Gemma 3 27B Preview
Google DeepMind🔓 Open Weights
76.2%
80.5%
75.8%
79.8%
74.2%
74.0%
88 tok/s
$0.25/$0.50
#36Llama 4 Scout
Meta AI🔓 Open Weights
76.0%
87.0%
77.5%
87.0%
80.0%
89.0%
140 tok/s
$0.20/$0.60
#37Llama 3.3 70B Instruct
Meta AIFlagship
75.8%
80.5%
75.5%
79.5%
73.5%
72.5%
90 tok/s
$0.30/$0.60
#38Mistral Large 2
Mistral AIFlagship
75.2%
79.8%
75.0%
78.5%
73.0%
72.0%
Open-Weights
#39Ollama Llama 3.3 (Q4_K_M)
Ollama (Inference Runtime)🔓 Open Weights
75.0%
42 tok/sOpen-Weights
#40Gemini 2.0 Flash Thinking
Google DeepMind★ Frontier SOTA
74.9%
87.0%
75.2%
88.5%
79.0%
89.0%
95 tok/s
$0.10/$0.40
#41Hunyuan-Large
Tencent🔓 Open Weights
74.8%
65 tok/s
$0.60/$1.80
#42DeepSeek V4-Flash
DeepSeek🔓 Open Weights
74.5%
84.0%
75.5%
86.0%
72.0%
62.0%
190 tok/s
$0.14/$0.28
#43Amazon Nova Pro
Amazon AWSFlagship
74.2%
120 tok/s
$0.80/$3.20
#44Sonar
Perplexity AICloud API
74.0%
78.0%
73.0%
76.5%
71.5%
73.0%
92 tok/s
$0.50/$0.50
#45Gemini 3.5 Flash-Lite
Google DeepMindFlagship
73.8%
83.5%
74.0%
85.2%
71.0%
61.5%
220 tok/s
$0.10/$0.40
#46QwQ 32B
Alibaba Cloud / Qwen🔓 Open Weights
73.8%
86.2%
74.8%
89.1%
75.2%
85.0%
45 tok/s
$0.20/$0.60
#47o1
OpenAI★ Frontier SOTA
73.5%
86.8%
74.0%
89.2%
77.0%
87.5%
45 tok/s
$15.00/$60.00
#48Yi-Lightning
01.AIFlagship
73.5%
110 tok/s
$0.14/$0.14
#49MiniMax-Text-01
MiniMaxFlagship
73.0%
75 tok/s
$0.20/$1.10
#50Gemini 2.0 Pro
Google DeepMind★ Frontier SOTA
72.8%
85.5%
73.5%
85.0%
80.1%
88.0%
60 tok/s
$1.50/$6.00
#51GLM-4-Plus
Zhipu AIFlagship
72.5%
65 tok/s
$1.40/$1.40
#52Mistral Large 3
Mistral AIFlagship
72.0%
85.0%
73.0%
84.0%
78.0%
88.0%
60 tok/s
$2.00/$6.00
#53ERNIE 4.0 Turbo
BaiduFlagship
71.8%
58 tok/s
$4.20/$8.40
#54Gemma 2 27B
Google DeepMind🔓 Open Weights
71.5%
74.5%
69.5%
73.0%
70.0%
71.0%
95 tok/s
$0.20/$0.40
#55Qwen 2.5 Max
Alibaba Cloud / QwenFlagship
71.4%
84.5%
71.5%
83.0%
76.5%
87.0%
58 tok/s
$1.60/$6.40
#56GPT-4.5
OpenAI★ Frontier SOTA
71.0%
85.0%
70.5%
80.0%
81.5%
91.5%
42 tok/s
$75.00/$150.00
#57Kimi k1.5
Moonshot AIFlagship
70.5%
48 tok/s
$1.00/$4.00
#58Claude 3.5 Sonnet
AnthropicFlagship
69.8%
85.0%
72.8%
78.4%
77.8%
89.5%
75 tok/s
$3.00/$15.00
#59Amazon Nova Lite
Amazon AWSFlagship
69.5%
200 tok/s
$0.06/$0.24
#60Reka Core
Reka AIFlagship
69.2%
55 tok/s
$3.00/$15.00
#61Solar Pro
UpstageFlagship
68.0%
90 tok/s
$0.25/$0.25
#62DeepSeek-V3
DeepSeek🔓 Open Weights
67.8%
81.9%
69.2%
78.9%
74.2%
87.5%
62 tok/s
$0.14/$0.28
#63Gemini 2.0 Flash
Google DeepMindFlagship
67.5%
82.0%
68.4%
79.1%
75.0%
88.7%
185 tok/s
$0.10/$0.40
#64DBRX Instruct
Databricks🔓 Open Weights
67.2%
70 tok/s
$0.60/$0.60
#65Qwen 2.5 Coder 32B
Alibaba Cloud / Qwen🔓 Open Weights
67.0%
74.0%
75 tok/s
$0.20/$0.60
#66Llama 3.1 405B
Meta AI🔓 Open Weights
66.8%
81.0%
67.0%
77.5%
74.0%
86.5%
32 tok/s
$2.00/$2.00
#67Gemma 2 9B
Google DeepMind🔓 Open Weights
66.8%
69.0%
64.0%
68.0%
65.5%
67.5%
130 tok/s
$0.10/$0.20
#68Snowflake Arctic
Snowflake🔓 Open Weights
66.5%
75 tok/s
$0.80/$0.80
#69GPT-4o
OpenAIFlagship
66.4%
81.5%
66.5%
76.8%
78.4%
90.8%
70 tok/s
$2.50/$10.00
#70Llama 3.3 70B
Meta AI🔓 Open Weights
65.2%
79.5%
65.1%
76.2%
73.0%
86.2%
72 tok/s
$0.40/$0.40
#71Gemini 1.5 Pro
Google DeepMindFlagship
64.1%
55 tok/s
$1.25/$5.00
#72Mistral Large 2
Mistral AIFlagship
63.4%
55 tok/s
$2.00/$6.00
#73Gemma 2 2B
Google DeepMind🔓 Open Weights
56.4%
58.0%
52.0%
54.5%
55.0%
60.0%
180 tok/s
$0.05/$0.10
#74GPT-4o mini
OpenAIFlagship
56.2%
125 tok/s
$0.15/$0.60
#75Sonar Large
Perplexity AIFlagship
85 tok/s
$1.00/$1.00
#76Codestral 25.01
Mistral AI🔓 Open Weights
90 tok/s
$0.20/$0.60
#77Claude 3 Opus
AnthropicFlagship
30 tok/s
$15.00/$75.00
#78Claude 3.5 Haiku
AnthropicFlagship
95 tok/s
$0.80/$4.00
#79Phi-4
Microsoft Research🔓 Open Weights
85 tok/s
$0.10/$0.20
#80Command R+
CohereFlagship
50 tok/s
$2.50/$10.00
#81Gemini 1.5 Flash
Google DeepMindFlagship
140 tok/s
$0.07/$0.30
#82GPT-4.5 (pure LLM)
OpenAICloud API
Open-Weights
#83r1 / r1-zero (single CoT)
DeepSeek🔓 Open Weights
Open-Weights
#84ARChitects (Kaggle 2024 winner)
OpenAICloud API
Open-Weights
#85o3-mini high (Epoch independent eval, Tiers 1–3, tools-on)
OpenAICloud API
Open-Weights
#86Goedel-Prover (Lean, 7/644, pass@512)
OpenAICloud API
Open-Weights
#87o3-mini (CodeSOTA MBPP pass@1)
OpenAICloud API
Open-Weights
#88DeepSeek-V3 (base, 0-shot)
DeepSeek🔓 Open Weights
Open-Weights
#89DeepSeek V3 (open weight)
DeepSeek🔓 Open Weights
Open-Weights
#90DeepSeek-V3 (CodeSOTA MBPP pass@1)
DeepSeek🔓 Open Weights
Open-Weights
#91o3-preview-low (CoT + search/synthesis)
OpenAICloud API
Open-Weights
#92o3 (AIME 2024, pass@1)
OpenAICloud API
Open-Weights
#93Jeremy Berman system
OpenAICloud API
Open-Weights
#94Qwen2.5-Coder-32B-Instruct (EvalPlus, greedy, HumanEval base pass@1)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#95Qwen2.5-Coder-32B-Instruct
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#96Qwen2.5-Coder-32B-Instruct (CodeSOTA MBPP pass@1)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#97Qwen2.5-Coder 32B (CodeSOTA MBPP pass@1)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#98best frontier model (Tiers 1–3, tools-on)
OpenAICloud API
Open-Weights
#99Grok Beta
xAICloud API
Open-Weights
#100DeepSeek-V3 (November 2024)
DeepSeek🔓 Open Weights
Open-Weights
#101DeepSeek-V2.5 (November 2024)
DeepSeek🔓 Open Weights
Open-Weights
#102Claude 3.5 Sonnet (Oct 2024; CodeSOTA MBPP pass@1)
AnthropicCloud API
Open-Weights
#103Claude 3.5 Sonnet (upgraded; Anthropic tools scaffold)
AnthropicCloud API
Open-Weights
#104OpenAI o1-mini
OpenAICloud API
Open-Weights
#105Gemini 1.5 Pro 002
Google DeepMindCloud API
Open-Weights
#106Qwen2.5 72B (base, 0-shot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#107o1-preview
OpenAICloud API
Open-Weights
#108GPT-4o (AIME 2024, pass@1)
OpenAICloud API
Open-Weights
#109o1 (AIME 2024, pass@1)
OpenAICloud API
Open-Weights
#110o1-preview (September 2024)
OpenAICloud API
Open-Weights
#111o1-mini (September 2024)
OpenAICloud API
Open-Weights
#112o1 (MATH-500)
OpenAICloud API
Open-Weights
#113O1 Preview (Sept 2024)
OpenAICloud API
Open-Weights
#114O1 Mini (Sept 2024)
OpenAICloud API
Open-Weights
#115GPT-4o (Agentless)
OpenAICloud API
Open-Weights
#116GPT-4o (August 2024)
OpenAICloud API
Open-Weights
#117Jamba-1.5-large (README Avg. snapshot)
AI21 LabsCloud API
Open-Weights
#118Gemini-1.5-pro (README Avg. snapshot)
Google DeepMindCloud API
Open-Weights
#119Qwen2.5-14B-Instruct-1M (README Avg. snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#120Qwen3-235B-A22B (README Avg. snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#121Qwen3-14B (README Avg. snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#122Jamba-1.5-mini (README Avg. snapshot)
AI21 LabsCloud API
Open-Weights
#123Qwen3-32B (README Avg. snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#124EXAONE-4.0-32B (README Avg. snapshot)
LG AI ResearchCloud API
Open-Weights
#125Qwen2.5-7B-Instruct-1M (README Avg. snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#126GPT-4-1106-preview (README Avg. snapshot)
OpenAICloud API
Open-Weights
#127Claude Fable 5 (LLM Stats snapshot)
AnthropicCloud API
Open-Weights
#128Claude Mythos Preview (LLM Stats snapshot)
AnthropicCloud API
Open-Weights
#129Claude Opus 4.8 (LLM Stats snapshot)
AnthropicCloud API
Open-Weights
#130Grok 4.5 (LLM Stats snapshot)
xAICloud API
Open-Weights
#131GPT-5.6 Sol (LLM Stats snapshot)
OpenAICloud API
Open-Weights
#132Claude Opus 4.7 (LLM Stats snapshot)
AnthropicCloud API
Open-Weights
#133GPT-5.6 Terra (LLM Stats snapshot)
OpenAICloud API
Open-Weights
#134Claude Sonnet 5 (LLM Stats snapshot)
AnthropicCloud API
Open-Weights
#135GPT-5.6 Luna (LLM Stats snapshot)
OpenAICloud API
Open-Weights
#136GLM-5.2 (LLM Stats snapshot)
Zhipu AICloud API
Open-Weights
#137Muse Spark 1.1 (LLM Stats snapshot)
Meta AI🔓 Open Weights
Open-Weights
#138Qwen3.7 Max (LLM Stats snapshot)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#139MiniMax M3 (LLM Stats snapshot)
MiniMaxCloud API
Open-Weights
#140Gemini 3.6 Flash (LLM Stats snapshot)
Google DeepMindCloud API
Open-Weights
#141Kimi K2.6 (LLM Stats snapshot)
OpenAICloud API
Open-Weights
#142GPT-5.5 (LLM Stats snapshot)
OpenAICloud API
Open-Weights
#143GLM-5.1 (LLM Stats snapshot)
Zhipu AICloud API
Open-Weights
#144Llama 3.1 405B (5-shot)
Meta AI🔓 Open Weights
Open-Weights
#145GPT-4o Mini (July 2024)
OpenAICloud API
Open-Weights
#146GPT-4o (paper era, Lean, ~1/640)
OpenAICloud API
Open-Weights
#147Claude 3.5 Sonnet (3-shot CoT)
AnthropicCloud API
Open-Weights
#148Claude 3.5 Sonnet (June 2024)
AnthropicCloud API
Open-Weights
#149DeepSeek-Coder-V2-Instruct
DeepSeek🔓 Open Weights
Open-Weights
#150DeepSeek-Coder-V2-Instruct (CodeSOTA MBPP pass@1)
DeepSeek🔓 Open Weights
Open-Weights
#151o1-mini
OpenAICloud API
Open-Weights
#152Claude 4.7 (High)
AnthropicCloud API
Open-Weights
#153GPT-5.5 Pro (High)
OpenAICloud API
Open-Weights
#154GPT-5.6 Sol (xHigh)
OpenAICloud API
Open-Weights
#155NVARC
OpenAICloud API
Open-Weights
#156Claude 4.7 (Max)
AnthropicCloud API
Open-Weights
#157GPT-5.5 (xHigh)
OpenAICloud API
Open-Weights
#158GPT-5.6 Sol (Max)
OpenAICloud API
Open-Weights
#159Anthropic Opus 4.6 (Max)
AnthropicCloud API
Open-Weights
#160Gemini 3.1 Pro (Preview)
Google DeepMindCloud API
Open-Weights
#161GPT-5.5 (High)
OpenAICloud API
Open-Weights
#162Claude Opus 4.8 (High)
AnthropicCloud API
Open-Weights
#163SWE-1.6 (FrontierCode 1.1 Main pass rate)
OpenAICloud API
Open-Weights
#164Kimi K2.7 Code (FrontierCode 1.1 Main pass rate)
Moonshot AICloud API
Open-Weights
#165SWE-1.7 (FrontierCode 1.1 Main pass rate)
OpenAICloud API
Open-Weights
#166GPT-5.5 (FrontierCode 1.1 Main pass rate)
OpenAICloud API
Open-Weights
#167Claude Opus 4.8 (FrontierCode 1.1 Main pass rate)
AnthropicCloud API
Open-Weights
#168Qwen3.7 Max (qwen3-7-max; closed)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#169Qwen3.7 Plus (qwen3-7-plus; closed)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#170GLM-4.7 (glm-4-7; open weight)
Zhipu AICloud API
Open-Weights
#171Qwen3.6-27B (open weight)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#172Qwen3.6-35B-A3B (open weight)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#173Gemini-2.5-Pro (paper Overall)
Google DeepMindCloud API
Open-Weights
#174GPT-5 (paper Overall)
OpenAICloud API
Open-Weights
#175Qwen3-235B-A22B-Thinking-2507 (paper Overall)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#176Qwen3-Next-80B-A3B-Thinking (paper Overall)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#177DeepSeek-R1-0528 (paper Overall)
DeepSeek🔓 Open Weights
Open-Weights
#178DeepSeek-R1 (paper Overall)
DeepSeek🔓 Open Weights
Open-Weights
#179Qwen3-30B-A3B-Thinking-2507 (paper Overall)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#180Claude-4-Sonnet (paper Overall)
AnthropicCloud API
Open-Weights
#181Gemini-2.5-Flash (paper Overall)
Google DeepMindCloud API
Open-Weights
#182MiniMax-M2 (paper Overall)
MiniMaxCloud API
Open-Weights
#183GPT-5.5 (xhigh, expected performance)
OpenAICloud API
Open-Weights
#184GPT-5.6-Sol (max, expected performance)
OpenAICloud API
Open-Weights
#185o4-mini (CodeSOTA MBPP pass@1)
OpenAICloud API
Open-Weights
#186Claude Opus 4 (CodeSOTA MBPP pass@1)
AnthropicCloud API
Open-Weights
#187GPT-4.1 (CodeSOTA MBPP pass@1)
OpenAICloud API
Open-Weights
#188Claude Sonnet 4 (CodeSOTA MBPP pass@1)
AnthropicCloud API
Open-Weights
#189GPT-5.6 Sol (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#190GPT-5.6 Terra (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#191Claude Opus 4.6 (LLM Stats MRCR v2 8-needle)
AnthropicCloud API
Open-Weights
#192GPT-5.5 (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#193Gemma 4 31B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#194Gemini 3.1 Flash-Lite (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#195Gemini 3.6 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#196Gemma 4 26B-A4B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#197Gemma 4 12B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#198GPT-5.6 Luna (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#199GPT-5.4 mini (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#200GPT-5.4 nano (LLM Stats MRCR v2 8-needle)
OpenAICloud API
Open-Weights
#201Gemini 3.5 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#202Gemini 3.1 Pro (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#203Gemini 3 Pro (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#204Gemma 4 E4B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#205Gemini 3 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#206Gemini 3.5 Flash-Lite (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#207Gemma 4 E2B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#208Gemini 2.5 Pro Preview 06-05 (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#209Gemma 3 27B (LLM Stats MRCR v2 8-needle)
Google DeepMindCloud API
Open-Weights
#210GPT-5 (Lean, ~42/660, pass@1, ~10-turn ReAct)
OpenAICloud API
Open-Weights
#211Gemini 3 Pro (Challenge Avg@3)
Google DeepMindCloud API
Open-Weights
#212GPT-5.5 high (Challenge Avg@3)
OpenAICloud API
Open-Weights
#213Qwen3 32B + SWE-agent (public, launch revision)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#214GPT-4o + SWE-agent (public, launch revision)
OpenAICloud API
Open-Weights
#215Claude Sonnet 4 + SWE-agent (public, launch revision)
AnthropicCloud API
Open-Weights
#216Claude Opus 4.1 + SWE-agent (public, launch revision)
AnthropicCloud API
Open-Weights
#217GPT-5 + SWE-agent (public, launch revision)
OpenAICloud API
Open-Weights
#218Claude 4.5 Opus (high reasoning; mini-SWE-agent 2.0.0)
AnthropicCloud API
Open-Weights
#219Gemini 3 Flash (high reasoning; mini-SWE-agent 2.0.0)
Google DeepMindCloud API
Open-Weights
#220MiniMax M2.5 (high reasoning; mini-SWE-agent 2.0.0)
MiniMaxCloud API
Open-Weights
#221Claude Opus 4.6 (mini-SWE-agent 2.0.0)
AnthropicCloud API
Open-Weights
#222GPT-5-2 Codex (mini-SWE-agent 2.0.0)
OpenAICloud API
Open-Weights
#223GLM-5 (high reasoning; mini-SWE-agent 2.0.0)
Zhipu AICloud API
Open-Weights
#224GPT-5-2 (high reasoning; mini-SWE-agent 2.0.0)
OpenAICloud API
Open-Weights
#225Claude 4.5 Sonnet (high reasoning; mini-SWE-agent 2.0.0)
AnthropicCloud API
Open-Weights
#226Kimi K2.5 (high reasoning; mini-SWE-agent 2.0.0)
OpenAICloud API
Open-Weights
#227DeepSeek V3.2 (high reasoning; mini-SWE-agent 2.0.0)
DeepSeek🔓 Open Weights
Open-Weights
#228Gemini 3 Pro (mini-SWE-agent 2.0.0)
Google DeepMindCloud API
Open-Weights
#229Claude 4.5 Haiku (high reasoning; mini-SWE-agent 2.0.0)
AnthropicCloud API
Open-Weights
#230GPT-5 Mini (mini-SWE-agent 2.0.0)
OpenAICloud API
Open-Weights
#231GPT-4o (198-task Diamond offline)
OpenAICloud API
Open-Weights
#232o1 (198-task Diamond offline)
OpenAICloud API
Open-Weights
#233Claude Code + Fable 5 (xhigh; TB 2.1)
AnthropicCloud API
Open-Weights
#234Codex + GPT-5.5 (xhigh; TB 2.1)
OpenAICloud API
Open-Weights
#235Terminus 2 + Fable 5 (high; TB 2.1)
AnthropicCloud API
Open-Weights
#236Cursor CLI + Grok 4.5 (high; TB 2.1)
xAICloud API
Open-Weights
#237Claude Code + Opus 4.8 (high; TB 2.1)
AnthropicCloud API
Open-Weights
#238Codex + GPT-5.6 Terra (max; TB 2.1)
OpenAICloud API
Open-Weights
#239Terminus 2 + GPT-5.5 (xhigh; TB 2.1)
OpenAICloud API
Open-Weights
#240mini-SWE-agent + Muse Spark 1.1 (xhigh; TB 2.1)
Meta AI🔓 Open Weights
Open-Weights
#241Codex + GPT-5.6 Luna (max; TB 2.1)
OpenAICloud API
Open-Weights
#242Claude Code + Sonnet 5 (high; TB 2.1)
AnthropicCloud API
Open-Weights
#243Terminus 2 + Gemini 3 Pro (high; TB 2.1)
Google DeepMindCloud API
Open-Weights
#244Claude Code + Opus 4.7 (max; TB 2.1)
AnthropicCloud API
Open-Weights
#245Terminus 2 + Opus 4.7 (max; TB 2.1)
AnthropicCloud API
Open-Weights
#246Gemini CLI + Gemini 3 Pro (high; TB 2.1)
Google DeepMindCloud API
Open-Weights
#247Gemini CLI + Gemini 3.1 Pro (high; TB 2.1)
Google DeepMindCloud API
Open-Weights
#248Terminus 2 + Gemini 3.1 Pro (high; TB 2.1)
Google DeepMindCloud API
Open-Weights
#249Claude Code + GLM-5.1 (max; TB 2.1)
Zhipu AICloud API
Open-Weights
#250GPT-4o (few-shot)
OpenAICloud API
Open-Weights
#251GPT-4o (full benchmark)
OpenAICloud API
Open-Weights
#252GPT-4 Turbo + SWE-agent
OpenAICloud API
Open-Weights
#253Gemini 1.5 Pro (3-shot CoT)
Google DeepMindCloud API
Open-Weights
#254Llama 3 70B (7-shot)
Meta AI🔓 Open Weights
Open-Weights
#255llama 3 70b-instruct
Meta AI🔓 Open Weights
Open-Weights
#256llama 3 8b-instruct
Meta AI🔓 Open Weights
Open-Weights
#257Llama 3 70B
Meta AI🔓 Open Weights
Open-Weights
#258CodeQwen1.5-7B-Chat
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#259CodeQwen1.5-7B base (Mercury-eval Overall Beyond, 5 samples)
Alibaba Cloud / Qwen🔓 Open Weights
Open-Weights
#260Llama 3 70B (5-shot)
Meta AI🔓 Open Weights
Open-Weights
#261GPT-4-Turbo (April 2024)
OpenAICloud API
Open-Weights
#262GPT-4-Turbo
OpenAICloud API
Open-Weights
#263Claude 3 Opus (3-shot CoT)
AnthropicCloud API
Open-Weights
#264Claude 3 Opus (BM25 retrieval)
AnthropicCloud API
Open-Weights
#265Claude 3 Opus (5-shot)
AnthropicCloud API
Open-Weights
#266StarCoder2-15B base (Mercury-eval Overall Beyond, 5 samples)
OpenAICloud API
Open-Weights
#267GPT-4 (paper Table 3 Average)
OpenAICloud API
Open-Weights
#268YaRN-Mistral (paper Table 3 Average)
Mistral AI🔓 Open Weights
Open-Weights
#269Kimi-Chat (paper Table 3 Average)
OpenAICloud API
Open-Weights
#270Claude 2 (paper Table 3 Average)
AnthropicCloud API
Open-Weights
#271GPT-4V (full benchmark)
OpenAICloud API
Open-Weights
#272UN codellama-70b-instruct
Meta AI🔓 Open Weights
Open-Weights
#273Gemini Ultra 1.0 (3-shot CoT)
Google DeepMindCloud API
Open-Weights
#274Gemini Ultra (10-shot decontaminated)
Google DeepMindCloud API
Open-Weights
#275PaLM 540B (translate-to-English CoT)
Google DeepMindCloud API
Open-Weights
#276GPT-4-Turbo (November 2023)
OpenAICloud API
Open-Weights
#277DeepSeek-Coder-33B base (Mercury-eval Overall Beyond, 5 samples)
DeepSeek🔓 Open Weights
Open-Weights
#278Claude 2 (BM25 retrieval)
AnthropicCloud API
Open-Weights
#279GPT-3.5-Turbo-16k (paper Table 3 OverAll)
OpenAICloud API
Open-Weights
#280Llama2-7B-chat-4k (paper Table 3 OverAll)
Meta AI🔓 Open Weights
Open-Weights
#281LongChat-v1.5-7B-32k (paper Table 3 OverAll)
OpenAICloud API
Open-Weights
#282XGen-7B-8k (paper Table 3 OverAll)
OpenAICloud API
Open-Weights
#283InternLM-7B-8k (paper Table 3 OverAll)
OpenAICloud API
Open-Weights
#284ChatGLM2-6B (paper Table 3 OverAll)
Zhipu AICloud API
Open-Weights
#285ChatGLM2-6B-32k (paper Table 3 OverAll)
Zhipu AICloud API
Open-Weights
#286Vicuna-v1.5-7B-16k (paper Table 3 OverAll)
OpenAICloud API
Open-Weights
#287UN codellama-34b-instruct
Meta AI🔓 Open Weights
Open-Weights
#288UN codellama-13b-instruct
Meta AI🔓 Open Weights
Open-Weights
#289CodeLlama-34B base (Mercury-eval Overall Beyond, 5 samples)
Meta AI🔓 Open Weights
Open-Weights
#290Llama 2 70B (5-shot)
Meta AI🔓 Open Weights
Open-Weights
#291T0pp (paper Table 3 Avg)
OpenAICloud API
Open-Weights
#292Flan-T5 (paper Table 3 Avg)
Google DeepMindCloud API
Open-Weights
#293Flan-UL2 (paper Table 3 Avg)
Google DeepMindCloud API
Open-Weights
#294DaVinci003 (paper Table 3 Avg)
OpenAICloud API
Open-Weights
#295ChatGPT (paper Table 3 Avg)
OpenAICloud API
Open-Weights
#296Claude (paper Table 3 Avg)
AnthropicCloud API
Open-Weights
#297GPT-4 (paper Table 3 Avg)
OpenAICloud API
Open-Weights
#298CPACE
OpenAICloud API
Open-Weights
#299GPT-4 (EvalPlus Table 3, greedy, HumanEval base pass@1)
OpenAICloud API
Open-Weights
#300GPT-4 (May 2023)
OpenAICloud API
Open-Weights
#301Chameleon (GPT-4, Text-GT, tools)
OpenAICloud API
Open-Weights
#302text-davinci-003
OpenAICloud API
Open-Weights
#303ChatGPT (gpt-3.5-turbo)
OpenAICloud API
Open-Weights
#304GPT-4
OpenAICloud API
Open-Weights
#305GPT-4 (3-shot CoT)
OpenAICloud API
Open-Weights
#306GPT-4 (5-shot CoT)
OpenAICloud API
Open-Weights
#307GPT-4 base (10-shot)
OpenAICloud API
Open-Weights
#308GPT-4 (full MATH)
OpenAICloud API
Open-Weights
#309GPT-4 (5-shot)
OpenAICloud API
Open-Weights
#310gpt-3.5-turbo
OpenAICloud API
Open-Weights
#311LLaMA-65B
Meta AI🔓 Open Weights
Open-Weights
#312SantaCoder-1.1B (MultiPL-HumanEval Python, pass@1)
OpenAICloud API
Open-Weights
#313PaLM 540B
Google DeepMindCloud API
Open-Weights
#314Codex (code-davinci-002)
OpenAICloud API
Open-Weights
#315Vega v2
OpenAICloud API
Open-Weights
#316GPT-3 + PromptPG (2-shot CoT, Text-GT)
OpenAICloud API
Open-Weights
#317InCoder-6.7B (MultiPL-HumanEval Python, pass@1)
Meta AI🔓 Open Weights
Open-Weights
#318code-davinci-002 (MultiPL-HumanEval Python, pass@1)
OpenAICloud API
Open-Weights
#319PaLM 540B (5-shot)
Google DeepMindCloud API
Open-Weights
#320Chinchilla 70B
Google DeepMindCloud API
Open-Weights
#321PaLM 540B (CoT + self-consistency)
Google DeepMindCloud API
Open-Weights
#322ST-MoE-32B
Google DeepMindCloud API
Open-Weights
#323AlphaCode 9B (test set, 10@100k, no clustering)
Google DeepMindCloud API
Open-Weights
#324AlphaCode 41B (test set, 10@100k, no clustering)
Google DeepMindCloud API
Open-Weights
#325AlphaCode 41B + clustering (test set, 10@100k)
Google DeepMindCloud API
Open-Weights
#326PaLM 540B (8-shot CoT)
Google DeepMindCloud API
Open-Weights
#327Naive (paper Table 2 Avg)
OpenAICloud API
Open-Weights
#328BART 256 (paper Table 2 Avg)
Meta AI🔓 Open Weights
Open-Weights
#329BART 512 (paper Table 2 Avg)
Meta AI🔓 Open Weights
Open-Weights
#330BART 1024 (paper Table 2 Avg)
Meta AI🔓 Open Weights
Open-Weights
#331LED 1024 (paper Table 2 Avg)
OpenAICloud API
Open-Weights
#332LED 4096 (paper Table 2 Avg)
OpenAICloud API
Open-Weights
#333LED 16384 (paper Table 2 Avg)
OpenAICloud API
Open-Weights
#334Gopher 280B
Google DeepMindCloud API
Open-Weights
#335KEAR ensemble
Microsoft Research🔓 Open Weights
Open-Weights
#336KEAR single model
Microsoft Research🔓 Open Weights
Open-Weights
#337GPT-3 175B (fine-tuned)
OpenAICloud API
Open-Weights
#338ERNIE 3.0
BaiduCloud API
Open-Weights
#339Codex-12B (original paper, 164-task HumanEval pass@1 estimator)
OpenAICloud API
Open-Weights
#340Random (paper Table 3 R-1)
OpenAICloud API
Open-Weights
#341Ext. Oracle (paper Table 3 R-1)
OpenAICloud API
Open-Weights
#342TextRank (paper Table 3 R-1)
OpenAICloud API
Open-Weights
#343PGNet (paper Table 3 R-1)
OpenAICloud API
Open-Weights
#344BART (paper Table 3 R-1)
Meta AI🔓 Open Weights
Open-Weights
#345HMNet* (paper Table 3 R-1)
Microsoft Research🔓 Open Weights
Open-Weights
#346PGNet (gold spans) (paper Table 3 R-1)
OpenAICloud API
Open-Weights
#347BART (gold spans) (paper Table 3 R-1)
Meta AI🔓 Open Weights
Open-Weights
#348HMNet (gold spans) (paper Table 3 R-1)
Microsoft Research🔓 Open Weights
Open-Weights
#349GPT-3 175B (full MATH)
OpenAICloud API
Open-Weights
#350Seq2Seq (CodeSearchNet summarization, six-language overall BLEU)
OpenAICloud API
Open-Weights
#351Transformer (CodeSearchNet summarization, six-language overall BLEU)
OpenAICloud API
Open-Weights
#352RoBERTa encoder (CodeSearchNet summarization, six-language overall BLEU)
Meta AI🔓 Open Weights
Open-Weights
#353CodeBERT encoder (CodeSearchNet summarization, six-language overall BLEU)
Microsoft Research🔓 Open Weights
Open-Weights
#354DeBERTa-xxlarge (1.5B)
Microsoft Research🔓 Open Weights
Open-Weights
#355DeBERTa (TuringNLRv4)
Microsoft Research🔓 Open Weights
Open-Weights
#356DeBERTa ensemble (TuringNLRv4)
Microsoft Research🔓 Open Weights
Open-Weights
#357DeBERTa 1.5B (single model)
Microsoft Research🔓 Open Weights
Open-Weights
#358GPT-3 175B
OpenAICloud API
Open-Weights
#359GPT-3 175B (few-shot)
OpenAICloud API
Open-Weights
#360ELECTRA-Large
Google DeepMindCloud API
Open-Weights
#361T5 (ensemble)
Google DeepMindCloud API
Open-Weights
#362T5-11B
Google DeepMindCloud API
Open-Weights
#363ALBERT (ensemble)
Google DeepMindCloud API
Open-Weights
#364RoBERTa-large
OpenAICloud API
Open-Weights
#365RoBERTa ensemble
OpenAICloud API
Open-Weights
#366RoBERTa
OpenAICloud API
Open-Weights
#367BERT-Large (fine-tuned)
Google DeepMindCloud API
Open-Weights
#368RoBERTa-large (fine-tuned)
OpenAICloud API
Open-Weights
#369XLNet (ensemble)
Google DeepMindCloud API
Open-Weights
#370BERT++
OpenAICloud API
Open-Weights
#371MT-DNN
Microsoft Research🔓 Open Weights
Open-Weights
#372BERT-Large
Google DeepMindCloud API
Open-Weights