AI Benchmark Leaderboard & Top Models Rankings
Verified evaluations across LiveBench contamination-free benchmarks, Artificial Analysis intelligence indices, and Autonomous Coding Agents as of September 2026.
LiveBench Contamination-Resistant Suite: Hard questions updated periodically with strict test-set isolation to prevent pre-training data memorization.
| Rank | Model & Lab | Global Avg | Reasoning | Coding | Math | Data | IF | Throughput | Pricing / 1M |
|---|---|---|---|---|---|---|---|---|---|
| #1 | Claude Fable 5.1 (Hybrid Reasoning) Anthropic★ Frontier SOTA | 84.7% | 89.7% | 86.4% | 88.5% | 82.1% | 81.3% | 54 tok/s | $5.00/$25.00 |
| #2 | GPT-6 Astra (max) OpenAI★ Frontier SOTA | 84.2% | 91.7% | 85.9% | 90.2% | 81.4% | 78.0% | 62 tok/s | $4.50/$20.00 |
| #3 | DeepSeek V4.1 Flash DeepSeek★ Frontier SOTA | 83.9% | 88.4% | 85.1% | 89.0% | 80.5% | 82.0% | 96 tok/s | $0.18/$0.55 |
| #4 | Claude Fable 5 Anthropic★ Frontier SOTA | 83.0% | 89.7% | 86.0% | 96.0% | 81.2% | 70.0% | 45 tok/s | $10.00/$50.00 |
| #5 | GPT-5.6 Sol OpenAI★ Frontier SOTA | 82.8% | 88.0% | 84.1% | 87.5% | 80.1% | 80.0% | 62 tok/s | $5.00/$30.00 |
| #6 | o3 OpenAI★ Frontier SOTA | 82.5% | 89.5% | 84.0% | 88.9% | 79.5% | 77.2% | 68 tok/s | $4.00/$18.00 |
| #7 | GPT-5 OpenAI★ Frontier SOTA | 82.0% | 94.8% | 84.5% | 94.0% | 88.5% | 95.5% | 85 tok/s | $2.50/$10.00 |
| #8 | Kimi K3 Moonshot AI🔓 Open Weights | 82.0% | 90.7% | 83.5% | 87.0% | 79.0% | 76.0% | 58 tok/s | $2.80/$14.00 |
| #9 | Claude 3.7 Sonnet Anthropic★ Frontier SOTA | 81.5% | 86.8% | 83.2% | 86.4% | 79.2% | 77.0% | 72 tok/s | $3.00/$15.00 |
| #10 | Grok 4 xAI★ Frontier SOTA | 81.2% | 94.5% | 84.0% | 95.0% | 87.0% | 94.0% | 75 tok/s | $5.00/$20.00 |
| #11 | Grok 3 xAI★ Frontier SOTA | 81.2% | 87.0% | 82.5% | 86.2% | 78.0% | 76.5% | 65 tok/s | $3.00/$15.00 |
| #12 | Gemini 2.5 Pro Google DeepMind★ Frontier SOTA | 81.0% | 86.5% | 82.0% | 86.0% | 78.5% | 77.0% | 85 tok/s | $1.25/$5.00 |
| #13 | DeepSeek-R1 DeepSeek🔓 Open Weights | 80.8% | 87.9% | 81.5% | 87.2% | 78.0% | 75.2% | 65 tok/s | $0.55/$2.19 |
| #14 | Sonar Reasoning Pro Perplexity AI★ Frontier SOTA | 80.5% | 87.0% | 81.2% | 86.0% | 78.0% | 75.2% | 78 tok/s | $2.00/$8.00 |
| #15 | Claude Sonnet 5 Anthropic★ Frontier SOTA | 80.4% | 93.0% | 83.2% | 91.5% | 86.5% | 94.8% | 95 tok/s | $3.00/$15.00 |
| #16 | GPT-5.5 Thinking OpenAI★ Frontier SOTA | 80.2% | 89.7% | 82.1% | 95.9% | 78.4% | 68.0% | 70 tok/s | $3.00/$15.00 |
| #17 | o3-mini OpenAI★ Frontier SOTA | 80.2% | 87.2% | 81.0% | 86.5% | 77.0% | 75.0% | 115 tok/s | $1.10/$4.40 |
| #18 | Claude Opus 5 Anthropic★ Frontier SOTA | 80.1% | 91.2% | 81.4% | 95.7% | 78.0% | 67.5% | 48 tok/s | $5.00/$25.00 |
| #19 | Ollama DeepSeek-R1 (Q4_K_M) Ollama (Inference Runtime)🔓 Open Weights | 80.0% | — | — | — | — | — | 24 tok/s | Open-Weights |
| #20 | Gemini 3 Pro Google DeepMind★ Frontier SOTA | 79.8% | 92.0% | 81.5% | 92.0% | 88.0% | 93.5% | 80 tok/s | $2.50/$10.00 |
| #21 | Gemini 2.5 Flash (Thinking) Google DeepMindCloud API | 79.8% | 85.2% | 80.5% | 84.8% | 76.5% | 77.0% | 145 tok/s | $0.07/$0.30 |
| #22 | QwQ-32B Preview Alibaba Cloud / QwenCloud API | 79.2% | 86.0% | 80.0% | 85.5% | 76.0% | 75.0% | — | Open-Weights |
| #23 | Gemini 3.7 Flash Google DeepMind★ Frontier SOTA | 78.8% | 87.8% | 78.9% | 93.5% | 76.5% | 65.9% | 182 tok/s | $0.75/$3.75 |
| #24 | Grok 4.6 xAI★ Frontier SOTA | 78.5% | 88.4% | 80.2% | 92.1% | 75.8% | 65.1% | 75 tok/s | $2.00/$6.00 |
| #25 | Claude 3.5 Sonnet (v2) AnthropicCloud API | 78.5% | 83.5% | 79.2% | 82.0% | 76.0% | 75.0% | — | Open-Weights |
| #26 | DeepSeek V4-Pro DeepSeek🔓 Open Weights | 77.9% | 87.1% | 81.0% | 91.4% | 75.2% | 64.8% | 65 tok/s | $0.66/$1.98 |
| #27 | Qwen3-235B Alibaba Cloud / Qwen🔓 Open Weights | 77.4% | 87.5% | 80.5% | 91.0% | 74.8% | 64.0% | 70 tok/s | $0.80/$2.40 |
| #28 | Gemini 3.6 Flash Google DeepMind★ Frontier SOTA | 77.2% | 86.5% | 77.8% | 91.2% | 75.0% | 64.5% | 175 tok/s | $0.75/$3.75 |
| #29 | Llama 4 Maverick Meta AI🔓 Open Weights | 77.1% | 88.2% | 79.0% | 89.5% | 82.0% | 90.0% | 82 tok/s | $0.50/$1.50 |
| #30 | GLM-5.2 Zhipu AI★ Frontier SOTA | 77.0% | 86.8% | 79.5% | 90.5% | 76.0% | 66.0% | 72 tok/s | $1.50/$6.00 |
| #31 | Sonar Pro Perplexity AIFlagship | 77.0% | 81.5% | 76.5% | 80.0% | 74.5% | 75.5% | 85 tok/s | $1.00/$1.00 |
| #32 | MiniMax M3 MiniMax★ Frontier SOTA | 76.8% | — | — | — | — | — | 88 tok/s | $0.30/$1.20 |
| #33 | QwQ Plus Alibaba Cloud / Qwen🔓 Open Weights | 76.5% | 88.0% | 79.0% | 92.5% | 74.0% | 63.0% | 52 tok/s | $0.80/$2.40 |
| #34 | Qwen 2.5 72B Instruct Alibaba Cloud / Qwen🔓 Open Weights | 76.5% | 81.0% | 77.0% | 81.5% | 74.0% | 72.0% | 82 tok/s | $0.35/$0.70 |
| #35 | Gemma 3 27B Preview Google DeepMind🔓 Open Weights | 76.2% | 80.5% | 75.8% | 79.8% | 74.2% | 74.0% | 88 tok/s | $0.25/$0.50 |
| #36 | Llama 4 Scout Meta AI🔓 Open Weights | 76.0% | 87.0% | 77.5% | 87.0% | 80.0% | 89.0% | 140 tok/s | $0.20/$0.60 |
| #37 | Llama 3.3 70B Instruct Meta AIFlagship | 75.8% | 80.5% | 75.5% | 79.5% | 73.5% | 72.5% | 90 tok/s | $0.30/$0.60 |
| #38 | Mistral Large 2 Mistral AIFlagship | 75.2% | 79.8% | 75.0% | 78.5% | 73.0% | 72.0% | — | Open-Weights |
| #39 | Ollama Llama 3.3 (Q4_K_M) Ollama (Inference Runtime)🔓 Open Weights | 75.0% | — | — | — | — | — | 42 tok/s | Open-Weights |
| #40 | Gemini 2.0 Flash Thinking Google DeepMind★ Frontier SOTA | 74.9% | 87.0% | 75.2% | 88.5% | 79.0% | 89.0% | 95 tok/s | $0.10/$0.40 |
| #41 | Hunyuan-Large Tencent🔓 Open Weights | 74.8% | — | — | — | — | — | 65 tok/s | $0.60/$1.80 |
| #42 | DeepSeek V4-Flash DeepSeek🔓 Open Weights | 74.5% | 84.0% | 75.5% | 86.0% | 72.0% | 62.0% | 190 tok/s | $0.14/$0.28 |
| #43 | Amazon Nova Pro Amazon AWSFlagship | 74.2% | — | — | — | — | — | 120 tok/s | $0.80/$3.20 |
| #44 | Sonar Perplexity AICloud API | 74.0% | 78.0% | 73.0% | 76.5% | 71.5% | 73.0% | 92 tok/s | $0.50/$0.50 |
| #45 | Gemini 3.5 Flash-Lite Google DeepMindFlagship | 73.8% | 83.5% | 74.0% | 85.2% | 71.0% | 61.5% | 220 tok/s | $0.10/$0.40 |
| #46 | QwQ 32B Alibaba Cloud / Qwen🔓 Open Weights | 73.8% | 86.2% | 74.8% | 89.1% | 75.2% | 85.0% | 45 tok/s | $0.20/$0.60 |
| #47 | o1 OpenAI★ Frontier SOTA | 73.5% | 86.8% | 74.0% | 89.2% | 77.0% | 87.5% | 45 tok/s | $15.00/$60.00 |
| #48 | Yi-Lightning 01.AIFlagship | 73.5% | — | — | — | — | — | 110 tok/s | $0.14/$0.14 |
| #49 | MiniMax-Text-01 MiniMaxFlagship | 73.0% | — | — | — | — | — | 75 tok/s | $0.20/$1.10 |
| #50 | Gemini 2.0 Pro Google DeepMind★ Frontier SOTA | 72.8% | 85.5% | 73.5% | 85.0% | 80.1% | 88.0% | 60 tok/s | $1.50/$6.00 |
| #51 | GLM-4-Plus Zhipu AIFlagship | 72.5% | — | — | — | — | — | 65 tok/s | $1.40/$1.40 |
| #52 | Mistral Large 3 Mistral AIFlagship | 72.0% | 85.0% | 73.0% | 84.0% | 78.0% | 88.0% | 60 tok/s | $2.00/$6.00 |
| #53 | ERNIE 4.0 Turbo BaiduFlagship | 71.8% | — | — | — | — | — | 58 tok/s | $4.20/$8.40 |
| #54 | Gemma 2 27B Google DeepMind🔓 Open Weights | 71.5% | 74.5% | 69.5% | 73.0% | 70.0% | 71.0% | 95 tok/s | $0.20/$0.40 |
| #55 | Qwen 2.5 Max Alibaba Cloud / QwenFlagship | 71.4% | 84.5% | 71.5% | 83.0% | 76.5% | 87.0% | 58 tok/s | $1.60/$6.40 |
| #56 | GPT-4.5 OpenAI★ Frontier SOTA | 71.0% | 85.0% | 70.5% | 80.0% | 81.5% | 91.5% | 42 tok/s | $75.00/$150.00 |
| #57 | Kimi k1.5 Moonshot AIFlagship | 70.5% | — | — | — | — | — | 48 tok/s | $1.00/$4.00 |
| #58 | Claude 3.5 Sonnet AnthropicFlagship | 69.8% | 85.0% | 72.8% | 78.4% | 77.8% | 89.5% | 75 tok/s | $3.00/$15.00 |
| #59 | Amazon Nova Lite Amazon AWSFlagship | 69.5% | — | — | — | — | — | 200 tok/s | $0.06/$0.24 |
| #60 | Reka Core Reka AIFlagship | 69.2% | — | — | — | — | — | 55 tok/s | $3.00/$15.00 |
| #61 | Solar Pro UpstageFlagship | 68.0% | — | — | — | — | — | 90 tok/s | $0.25/$0.25 |
| #62 | DeepSeek-V3 DeepSeek🔓 Open Weights | 67.8% | 81.9% | 69.2% | 78.9% | 74.2% | 87.5% | 62 tok/s | $0.14/$0.28 |
| #63 | Gemini 2.0 Flash Google DeepMindFlagship | 67.5% | 82.0% | 68.4% | 79.1% | 75.0% | 88.7% | 185 tok/s | $0.10/$0.40 |
| #64 | DBRX Instruct Databricks🔓 Open Weights | 67.2% | — | — | — | — | — | 70 tok/s | $0.60/$0.60 |
| #65 | Qwen 2.5 Coder 32B Alibaba Cloud / Qwen🔓 Open Weights | 67.0% | — | 74.0% | — | — | — | 75 tok/s | $0.20/$0.60 |
| #66 | Llama 3.1 405B Meta AI🔓 Open Weights | 66.8% | 81.0% | 67.0% | 77.5% | 74.0% | 86.5% | 32 tok/s | $2.00/$2.00 |
| #67 | Gemma 2 9B Google DeepMind🔓 Open Weights | 66.8% | 69.0% | 64.0% | 68.0% | 65.5% | 67.5% | 130 tok/s | $0.10/$0.20 |
| #68 | Snowflake Arctic Snowflake🔓 Open Weights | 66.5% | — | — | — | — | — | 75 tok/s | $0.80/$0.80 |
| #69 | GPT-4o OpenAIFlagship | 66.4% | 81.5% | 66.5% | 76.8% | 78.4% | 90.8% | 70 tok/s | $2.50/$10.00 |
| #70 | Llama 3.3 70B Meta AI🔓 Open Weights | 65.2% | 79.5% | 65.1% | 76.2% | 73.0% | 86.2% | 72 tok/s | $0.40/$0.40 |
| #71 | Gemini 1.5 Pro Google DeepMindFlagship | 64.1% | — | — | — | — | — | 55 tok/s | $1.25/$5.00 |
| #72 | Mistral Large 2 Mistral AIFlagship | 63.4% | — | — | — | — | — | 55 tok/s | $2.00/$6.00 |
| #73 | Gemma 2 2B Google DeepMind🔓 Open Weights | 56.4% | 58.0% | 52.0% | 54.5% | 55.0% | 60.0% | 180 tok/s | $0.05/$0.10 |
| #74 | GPT-4o mini OpenAIFlagship | 56.2% | — | — | — | — | — | 125 tok/s | $0.15/$0.60 |
| #75 | Sonar Large Perplexity AIFlagship | — | — | — | — | — | — | 85 tok/s | $1.00/$1.00 |
| #76 | Codestral 25.01 Mistral AI🔓 Open Weights | — | — | — | — | — | — | 90 tok/s | $0.20/$0.60 |
| #77 | Claude 3 Opus AnthropicFlagship | — | — | — | — | — | — | 30 tok/s | $15.00/$75.00 |
| #78 | Claude 3.5 Haiku AnthropicFlagship | — | — | — | — | — | — | 95 tok/s | $0.80/$4.00 |
| #79 | Phi-4 Microsoft Research🔓 Open Weights | — | — | — | — | — | — | 85 tok/s | $0.10/$0.20 |
| #80 | Command R+ CohereFlagship | — | — | — | — | — | — | 50 tok/s | $2.50/$10.00 |
| #81 | Gemini 1.5 Flash Google DeepMindFlagship | — | — | — | — | — | — | 140 tok/s | $0.07/$0.30 |
| #82 | GPT-4.5 (pure LLM) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #83 | r1 / r1-zero (single CoT) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #84 | ARChitects (Kaggle 2024 winner) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #85 | o3-mini high (Epoch independent eval, Tiers 1–3, tools-on) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #86 | Goedel-Prover (Lean, 7/644, pass@512) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #87 | o3-mini (CodeSOTA MBPP pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #88 | DeepSeek-V3 (base, 0-shot) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #89 | DeepSeek V3 (open weight) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #90 | DeepSeek-V3 (CodeSOTA MBPP pass@1) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #91 | o3-preview-low (CoT + search/synthesis) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #92 | o3 (AIME 2024, pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #93 | Jeremy Berman system OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #94 | Qwen2.5-Coder-32B-Instruct (EvalPlus, greedy, HumanEval base pass@1) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #95 | Qwen2.5-Coder-32B-Instruct Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #96 | Qwen2.5-Coder-32B-Instruct (CodeSOTA MBPP pass@1) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #97 | Qwen2.5-Coder 32B (CodeSOTA MBPP pass@1) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #98 | best frontier model (Tiers 1–3, tools-on) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #99 | Grok Beta xAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #100 | DeepSeek-V3 (November 2024) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #101 | DeepSeek-V2.5 (November 2024) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #102 | Claude 3.5 Sonnet (Oct 2024; CodeSOTA MBPP pass@1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #103 | Claude 3.5 Sonnet (upgraded; Anthropic tools scaffold) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #104 | OpenAI o1-mini OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #105 | Gemini 1.5 Pro 002 Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #106 | Qwen2.5 72B (base, 0-shot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #107 | o1-preview OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #108 | GPT-4o (AIME 2024, pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #109 | o1 (AIME 2024, pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #110 | o1-preview (September 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #111 | o1-mini (September 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #112 | o1 (MATH-500) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #113 | O1 Preview (Sept 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #114 | O1 Mini (Sept 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #115 | GPT-4o (Agentless) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #116 | GPT-4o (August 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #117 | Jamba-1.5-large (README Avg. snapshot) AI21 LabsCloud API | — | — | — | — | — | — | — | Open-Weights |
| #118 | Gemini-1.5-pro (README Avg. snapshot) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #119 | Qwen2.5-14B-Instruct-1M (README Avg. snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #120 | Qwen3-235B-A22B (README Avg. snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #121 | Qwen3-14B (README Avg. snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #122 | Jamba-1.5-mini (README Avg. snapshot) AI21 LabsCloud API | — | — | — | — | — | — | — | Open-Weights |
| #123 | Qwen3-32B (README Avg. snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #124 | EXAONE-4.0-32B (README Avg. snapshot) LG AI ResearchCloud API | — | — | — | — | — | — | — | Open-Weights |
| #125 | Qwen2.5-7B-Instruct-1M (README Avg. snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #126 | GPT-4-1106-preview (README Avg. snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #127 | Claude Fable 5 (LLM Stats snapshot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #128 | Claude Mythos Preview (LLM Stats snapshot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #129 | Claude Opus 4.8 (LLM Stats snapshot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #130 | Grok 4.5 (LLM Stats snapshot) xAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #131 | GPT-5.6 Sol (LLM Stats snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #132 | Claude Opus 4.7 (LLM Stats snapshot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #133 | GPT-5.6 Terra (LLM Stats snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #134 | Claude Sonnet 5 (LLM Stats snapshot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #135 | GPT-5.6 Luna (LLM Stats snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #136 | GLM-5.2 (LLM Stats snapshot) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #137 | Muse Spark 1.1 (LLM Stats snapshot) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #138 | Qwen3.7 Max (LLM Stats snapshot) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #139 | MiniMax M3 (LLM Stats snapshot) MiniMaxCloud API | — | — | — | — | — | — | — | Open-Weights |
| #140 | Gemini 3.6 Flash (LLM Stats snapshot) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #141 | Kimi K2.6 (LLM Stats snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #142 | GPT-5.5 (LLM Stats snapshot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #143 | GLM-5.1 (LLM Stats snapshot) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #144 | Llama 3.1 405B (5-shot) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #145 | GPT-4o Mini (July 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #146 | GPT-4o (paper era, Lean, ~1/640) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #147 | Claude 3.5 Sonnet (3-shot CoT) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #148 | Claude 3.5 Sonnet (June 2024) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #149 | DeepSeek-Coder-V2-Instruct DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #150 | DeepSeek-Coder-V2-Instruct (CodeSOTA MBPP pass@1) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #151 | o1-mini OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #152 | Claude 4.7 (High) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #153 | GPT-5.5 Pro (High) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #154 | GPT-5.6 Sol (xHigh) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #155 | NVARC OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #156 | Claude 4.7 (Max) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #157 | GPT-5.5 (xHigh) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #158 | GPT-5.6 Sol (Max) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #159 | Anthropic Opus 4.6 (Max) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #160 | Gemini 3.1 Pro (Preview) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #161 | GPT-5.5 (High) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #162 | Claude Opus 4.8 (High) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #163 | SWE-1.6 (FrontierCode 1.1 Main pass rate) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #164 | Kimi K2.7 Code (FrontierCode 1.1 Main pass rate) Moonshot AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #165 | SWE-1.7 (FrontierCode 1.1 Main pass rate) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #166 | GPT-5.5 (FrontierCode 1.1 Main pass rate) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #167 | Claude Opus 4.8 (FrontierCode 1.1 Main pass rate) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #168 | Qwen3.7 Max (qwen3-7-max; closed) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #169 | Qwen3.7 Plus (qwen3-7-plus; closed) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #170 | GLM-4.7 (glm-4-7; open weight) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #171 | Qwen3.6-27B (open weight) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #172 | Qwen3.6-35B-A3B (open weight) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #173 | Gemini-2.5-Pro (paper Overall) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #174 | GPT-5 (paper Overall) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #175 | Qwen3-235B-A22B-Thinking-2507 (paper Overall) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #176 | Qwen3-Next-80B-A3B-Thinking (paper Overall) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #177 | DeepSeek-R1-0528 (paper Overall) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #178 | DeepSeek-R1 (paper Overall) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #179 | Qwen3-30B-A3B-Thinking-2507 (paper Overall) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #180 | Claude-4-Sonnet (paper Overall) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #181 | Gemini-2.5-Flash (paper Overall) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #182 | MiniMax-M2 (paper Overall) MiniMaxCloud API | — | — | — | — | — | — | — | Open-Weights |
| #183 | GPT-5.5 (xhigh, expected performance) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #184 | GPT-5.6-Sol (max, expected performance) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #185 | o4-mini (CodeSOTA MBPP pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #186 | Claude Opus 4 (CodeSOTA MBPP pass@1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #187 | GPT-4.1 (CodeSOTA MBPP pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #188 | Claude Sonnet 4 (CodeSOTA MBPP pass@1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #189 | GPT-5.6 Sol (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #190 | GPT-5.6 Terra (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #191 | Claude Opus 4.6 (LLM Stats MRCR v2 8-needle) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #192 | GPT-5.5 (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #193 | Gemma 4 31B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #194 | Gemini 3.1 Flash-Lite (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #195 | Gemini 3.6 Flash (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #196 | Gemma 4 26B-A4B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #197 | Gemma 4 12B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #198 | GPT-5.6 Luna (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #199 | GPT-5.4 mini (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #200 | GPT-5.4 nano (LLM Stats MRCR v2 8-needle) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #201 | Gemini 3.5 Flash (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #202 | Gemini 3.1 Pro (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #203 | Gemini 3 Pro (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #204 | Gemma 4 E4B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #205 | Gemini 3 Flash (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #206 | Gemini 3.5 Flash-Lite (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #207 | Gemma 4 E2B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #208 | Gemini 2.5 Pro Preview 06-05 (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #209 | Gemma 3 27B (LLM Stats MRCR v2 8-needle) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #210 | GPT-5 (Lean, ~42/660, pass@1, ~10-turn ReAct) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #211 | Gemini 3 Pro (Challenge Avg@3) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #212 | GPT-5.5 high (Challenge Avg@3) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #213 | Qwen3 32B + SWE-agent (public, launch revision) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #214 | GPT-4o + SWE-agent (public, launch revision) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #215 | Claude Sonnet 4 + SWE-agent (public, launch revision) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #216 | Claude Opus 4.1 + SWE-agent (public, launch revision) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #217 | GPT-5 + SWE-agent (public, launch revision) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #218 | Claude 4.5 Opus (high reasoning; mini-SWE-agent 2.0.0) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #219 | Gemini 3 Flash (high reasoning; mini-SWE-agent 2.0.0) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #220 | MiniMax M2.5 (high reasoning; mini-SWE-agent 2.0.0) MiniMaxCloud API | — | — | — | — | — | — | — | Open-Weights |
| #221 | Claude Opus 4.6 (mini-SWE-agent 2.0.0) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #222 | GPT-5-2 Codex (mini-SWE-agent 2.0.0) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #223 | GLM-5 (high reasoning; mini-SWE-agent 2.0.0) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #224 | GPT-5-2 (high reasoning; mini-SWE-agent 2.0.0) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #225 | Claude 4.5 Sonnet (high reasoning; mini-SWE-agent 2.0.0) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #226 | Kimi K2.5 (high reasoning; mini-SWE-agent 2.0.0) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #227 | DeepSeek V3.2 (high reasoning; mini-SWE-agent 2.0.0) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #228 | Gemini 3 Pro (mini-SWE-agent 2.0.0) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #229 | Claude 4.5 Haiku (high reasoning; mini-SWE-agent 2.0.0) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #230 | GPT-5 Mini (mini-SWE-agent 2.0.0) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #231 | GPT-4o (198-task Diamond offline) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #232 | o1 (198-task Diamond offline) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #233 | Claude Code + Fable 5 (xhigh; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #234 | Codex + GPT-5.5 (xhigh; TB 2.1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #235 | Terminus 2 + Fable 5 (high; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #236 | Cursor CLI + Grok 4.5 (high; TB 2.1) xAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #237 | Claude Code + Opus 4.8 (high; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #238 | Codex + GPT-5.6 Terra (max; TB 2.1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #239 | Terminus 2 + GPT-5.5 (xhigh; TB 2.1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #240 | mini-SWE-agent + Muse Spark 1.1 (xhigh; TB 2.1) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #241 | Codex + GPT-5.6 Luna (max; TB 2.1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #242 | Claude Code + Sonnet 5 (high; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #243 | Terminus 2 + Gemini 3 Pro (high; TB 2.1) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #244 | Claude Code + Opus 4.7 (max; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #245 | Terminus 2 + Opus 4.7 (max; TB 2.1) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #246 | Gemini CLI + Gemini 3 Pro (high; TB 2.1) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #247 | Gemini CLI + Gemini 3.1 Pro (high; TB 2.1) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #248 | Terminus 2 + Gemini 3.1 Pro (high; TB 2.1) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #249 | Claude Code + GLM-5.1 (max; TB 2.1) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #250 | GPT-4o (few-shot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #251 | GPT-4o (full benchmark) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #252 | GPT-4 Turbo + SWE-agent OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #253 | Gemini 1.5 Pro (3-shot CoT) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #254 | Llama 3 70B (7-shot) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #255 | llama 3 70b-instruct Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #256 | llama 3 8b-instruct Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #257 | Llama 3 70B Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #258 | CodeQwen1.5-7B-Chat Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #259 | CodeQwen1.5-7B base (Mercury-eval Overall Beyond, 5 samples) Alibaba Cloud / Qwen🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #260 | Llama 3 70B (5-shot) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #261 | GPT-4-Turbo (April 2024) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #262 | GPT-4-Turbo OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #263 | Claude 3 Opus (3-shot CoT) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #264 | Claude 3 Opus (BM25 retrieval) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #265 | Claude 3 Opus (5-shot) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #266 | StarCoder2-15B base (Mercury-eval Overall Beyond, 5 samples) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #267 | GPT-4 (paper Table 3 Average) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #268 | YaRN-Mistral (paper Table 3 Average) Mistral AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #269 | Kimi-Chat (paper Table 3 Average) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #270 | Claude 2 (paper Table 3 Average) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #271 | GPT-4V (full benchmark) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #272 | UN codellama-70b-instruct Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #273 | Gemini Ultra 1.0 (3-shot CoT) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #274 | Gemini Ultra (10-shot decontaminated) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #275 | PaLM 540B (translate-to-English CoT) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #276 | GPT-4-Turbo (November 2023) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #277 | DeepSeek-Coder-33B base (Mercury-eval Overall Beyond, 5 samples) DeepSeek🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #278 | Claude 2 (BM25 retrieval) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #279 | GPT-3.5-Turbo-16k (paper Table 3 OverAll) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #280 | Llama2-7B-chat-4k (paper Table 3 OverAll) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #281 | LongChat-v1.5-7B-32k (paper Table 3 OverAll) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #282 | XGen-7B-8k (paper Table 3 OverAll) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #283 | InternLM-7B-8k (paper Table 3 OverAll) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #284 | ChatGLM2-6B (paper Table 3 OverAll) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #285 | ChatGLM2-6B-32k (paper Table 3 OverAll) Zhipu AICloud API | — | — | — | — | — | — | — | Open-Weights |
| #286 | Vicuna-v1.5-7B-16k (paper Table 3 OverAll) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #287 | UN codellama-34b-instruct Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #288 | UN codellama-13b-instruct Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #289 | CodeLlama-34B base (Mercury-eval Overall Beyond, 5 samples) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #290 | Llama 2 70B (5-shot) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #291 | T0pp (paper Table 3 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #292 | Flan-T5 (paper Table 3 Avg) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #293 | Flan-UL2 (paper Table 3 Avg) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #294 | DaVinci003 (paper Table 3 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #295 | ChatGPT (paper Table 3 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #296 | Claude (paper Table 3 Avg) AnthropicCloud API | — | — | — | — | — | — | — | Open-Weights |
| #297 | GPT-4 (paper Table 3 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #298 | CPACE OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #299 | GPT-4 (EvalPlus Table 3, greedy, HumanEval base pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #300 | GPT-4 (May 2023) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #301 | Chameleon (GPT-4, Text-GT, tools) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #302 | text-davinci-003 OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #303 | ChatGPT (gpt-3.5-turbo) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #304 | GPT-4 OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #305 | GPT-4 (3-shot CoT) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #306 | GPT-4 (5-shot CoT) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #307 | GPT-4 base (10-shot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #308 | GPT-4 (full MATH) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #309 | GPT-4 (5-shot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #310 | gpt-3.5-turbo OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #311 | LLaMA-65B Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #312 | SantaCoder-1.1B (MultiPL-HumanEval Python, pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #313 | PaLM 540B Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #314 | Codex (code-davinci-002) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #315 | Vega v2 OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #316 | GPT-3 + PromptPG (2-shot CoT, Text-GT) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #317 | InCoder-6.7B (MultiPL-HumanEval Python, pass@1) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #318 | code-davinci-002 (MultiPL-HumanEval Python, pass@1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #319 | PaLM 540B (5-shot) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #320 | Chinchilla 70B Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #321 | PaLM 540B (CoT + self-consistency) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #322 | ST-MoE-32B Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #323 | AlphaCode 9B (test set, 10@100k, no clustering) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #324 | AlphaCode 41B (test set, 10@100k, no clustering) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #325 | AlphaCode 41B + clustering (test set, 10@100k) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #326 | PaLM 540B (8-shot CoT) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #327 | Naive (paper Table 2 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #328 | BART 256 (paper Table 2 Avg) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #329 | BART 512 (paper Table 2 Avg) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #330 | BART 1024 (paper Table 2 Avg) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #331 | LED 1024 (paper Table 2 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #332 | LED 4096 (paper Table 2 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #333 | LED 16384 (paper Table 2 Avg) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #334 | Gopher 280B Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #335 | KEAR ensemble Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #336 | KEAR single model Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #337 | GPT-3 175B (fine-tuned) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #338 | ERNIE 3.0 BaiduCloud API | — | — | — | — | — | — | — | Open-Weights |
| #339 | Codex-12B (original paper, 164-task HumanEval pass@1 estimator) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #340 | Random (paper Table 3 R-1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #341 | Ext. Oracle (paper Table 3 R-1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #342 | TextRank (paper Table 3 R-1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #343 | PGNet (paper Table 3 R-1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #344 | BART (paper Table 3 R-1) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #345 | HMNet* (paper Table 3 R-1) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #346 | PGNet (gold spans) (paper Table 3 R-1) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #347 | BART (gold spans) (paper Table 3 R-1) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #348 | HMNet (gold spans) (paper Table 3 R-1) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #349 | GPT-3 175B (full MATH) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #350 | Seq2Seq (CodeSearchNet summarization, six-language overall BLEU) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #351 | Transformer (CodeSearchNet summarization, six-language overall BLEU) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #352 | RoBERTa encoder (CodeSearchNet summarization, six-language overall BLEU) Meta AI🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #353 | CodeBERT encoder (CodeSearchNet summarization, six-language overall BLEU) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #354 | DeBERTa-xxlarge (1.5B) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #355 | DeBERTa (TuringNLRv4) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #356 | DeBERTa ensemble (TuringNLRv4) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #357 | DeBERTa 1.5B (single model) Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #358 | GPT-3 175B OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #359 | GPT-3 175B (few-shot) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #360 | ELECTRA-Large Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #361 | T5 (ensemble) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #362 | T5-11B Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #363 | ALBERT (ensemble) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #364 | RoBERTa-large OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #365 | RoBERTa ensemble OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #366 | RoBERTa OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #367 | BERT-Large (fine-tuned) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #368 | RoBERTa-large (fine-tuned) OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #369 | XLNet (ensemble) Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |
| #370 | BERT++ OpenAICloud API | — | — | — | — | — | — | — | Open-Weights |
| #371 | MT-DNN Microsoft Research🔓 Open Weights | — | — | — | — | — | — | — | Open-Weights |
| #372 | BERT-Large Google DeepMindCloud API | — | — | — | — | — | — | — | Open-Weights |