TrustTheBench Platform Sitemap & Directory Index
Comprehensive visual index of verified AI benchmark suites, foundation models, research labs, capability insights, and forensic evaluation datasets. All links are crawlable and refreshed hourly.
Core Platform Hubs & Interactive Tools
Trust Matrix (Home)
Platform CoreLive visual matrix of benchmark saturation, contamination risk, and verified rankings.
AI Benchmark Leaderboard
LeaderboardCross-model leaderboard tracking LiveBench, SWE-bench, and Artificial Analysis indices.
Autonomous Coding Agents
AgentsVerified coding agent rankings across SWE-bench Verified and repository-scale software tasks.
All Benchmarks Directory
DirectoryDirectory of trusted AI benchmarks with contamination audits and saturation status.
Top AI Models Directory
DirectoryComprehensive database of foundation models, pricing, speed (tokens/sec), and scores.
AI Research Labs & Providers
DirectoryDirectory of global AI organizations: OpenAI, Anthropic, Google DeepMind, DeepSeek, and Meta.
Benchmark Advisor
Interactive ToolInteractive evaluation advisor recommending non-saturated benchmark suites by model scale.
Capability Insights & Analysis
ResearchDeep research on reasoning, coding, math, test-time compute, and cost-efficiency frontiers.
Capability Domains & Research Insights
Reasoning & Logic Insights
ReasoningEvaluating multi-step deduction, GPQA Diamond, ARC-AGI-2, and test-time compute scaling.
Coding & Software Engineering Insights
CodingSWE-bench Verified, LiveCodeBench, and autonomous code synthesis benchmarks.
Mathematics & Olympiad Insights
MathematicsFrontierMath, AIME 2025, and formal mathematical proof synthesis performance.
Data Analysis & Scientific Insights
Data AnalysisScientific tabular comprehension, chart reasoning, and quantitative analytics benchmarks.
Complex Instruction Following Insights
InstructionIFEval, multi-turn system constraint adherence, and negative formatting compliance.
Speed, Latency & Pricing Economics
EconomicsTokens per second throughput, time-to-first-token (TTFT), and API cost per million tokens.
AI Benchmark Suites Directory
QuALITY
nearing-saturationMultiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.
MBPP+
nearing-saturationA cleaned 378-task MBPP evaluation with roughly 35 times more tests for stricter Python functional correctness.
Mercury
deprecatedPython code-synthesis benchmark that rewards both functional correctness and runtime efficiency against real solution distributions.
MGSM
nearing-saturationA human-translated 250-problem GSM8K subset for comparing grade-school mathematical reasoning across ten languages.
MMLU
saturated57-subject multiple-choice exam spanning STEM, humanities, and social sciences — the defining knowledge benchmark of the 2020–2024 era.
MMLU-Pro
nearing-saturationMMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.
MMLU-Redux
saturatedA manual error audit of MMLU: experts re-annotated thousands of questions and found ~6.5% are broken — the evidence base for MMLU's label-noise ceiling.
MT-Bench
nearing-saturation80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1–10 scale.
MultiPL-E
deprecatedParallel HumanEval and MBPP translations for comparing code generation across many programming languages.
Needle-in-a-Haystack
saturatedSynthetic retrieval stress test that hides facts at different depths and context lengths, then asks the model to recover them.
OlympiadBench
activeBilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.
Omni-MATH
deprecatedA 4,428-problem Olympiad mathematics suite spanning more than 33 subdomains and ten annotated difficulty levels.
OpenAI-MRCR
activeMRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.
PutnamBench
activeFormal theorem-proving on Putnam problems — models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.
QASPER
nearing-saturationScientific-paper QA over full NLP papers, with abstractive, extractive, yes/no, and unanswerable answers plus supporting evidence.
QMSum
nearing-saturationQuery-focused meeting summarization over long transcripts from academic, product, and committee meetings.
MBPP
saturatedMostly Basic Programming Problems: 974 crowd-sourced Python synthesis tasks intended for entry-level programmers.
RULER
nearing-saturationConfigurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.
SCROLLS
nearing-saturationSeven-task long-sequence suite covering summarization, QA, and NLI over naturally long English texts.
SimpleQA
active4,326 short fact-seeking questions chosen to stump GPT-4 — a hallucination test graded correct / incorrect / not-attempted.
Soohak
activeFresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.
StructFlowBench
nearing-saturation155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).
SuperGLUE
saturatedEight English NLU tasks spanning question answering, inference, causal reasoning, word sense, and coreference.
SWE-bench
deprecatedReal GitHub issues from 12 Python repositories, solved by generating patches that must pass repository tests.
SWE-bench Pro
activeLong-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.
SWE-bench Verified
saturatedA 500-task, engineer-reviewed subset of SWE-bench that became the standard repository-level coding-agent test.
SWE-Lancer
deprecatedCurrent 198-task offline patch benchmark derived from paid Upwork work, with historical mixed implementation and manager-task variants.
TabMWP
saturatedGrade-level math word problems that require combining a question with structured, textual, or image-rendered tables.
Terminal-Bench
activeEighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.
TruthfulQA
saturated817 adversarial questions probing whether models repeat common human misconceptions — famous for finding that bigger models were often less truthful.
WinoGrande
saturatedBinary fill-in-the-blank commonsense problems, scaled from Winograd schemas and adversarially filtered to reduce dataset shortcuts.
ZeroSCROLLS
nearing-saturationZero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.
GLUE
saturatedNine-task English NLU suite combining acceptability, sentiment, similarity, paraphrase, inference and coreference into one score.
AGIEval
saturatedStandardized human exams — SATs, LSATs, China's Gaokao and civil-service tests — repurposed to grade models against real human test-takers.
AIME
nearing-saturationThe MAA's competition exam — 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.
ARC-AGI
saturatedColored-grid abstraction tasks that test few-shot rule induction, not memorized facts, language fluency, or school math.
ARC-AGI-2
nearing-saturationThe harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.
ARC-AGI-3
activeInteractive ARC environments that score agents by human-relative action efficiency, planning, exploration, and adaptation.
BIG-Bench
deprecated204-task community-built mega-suite spanning reasoning, knowledge, math, code and social bias — the sprawling proving ground BBH was distilled from.
BIG-Bench Hard
saturated23 BIG-Bench tasks where models trailed humans — the benchmark that proved chain-of-thought prompting works, then fell to the reasoning-model era.
CodeContests
deprecatedDeepMind's competitive-programming corpus with temporally split problems, human submissions, and generated correctness tests.
CodeXGLUE
deprecatedA broad code-intelligence suite spanning understanding, retrieval, completion, translation, repair, generation, and summarization.
CommonsenseQA
saturatedFive-choice questions derived from ConceptNet relations, designed to require everyday knowledge beyond a supplied passage.
ComplexBench
nearing-saturationBilingual benchmark of how models follow instructions when several constraints must hold at once — where stacking constraints is not linear.
CyberSecEval
deprecatedMeta's evolving suite for insecure code, cyberattack assistance, prompt injection, exploitation, SOC analysis, and autopatching.
FireBench
activeEvaluates instruction following for enterprise/API pipelines — exact output format, strict step ordering, ranking, and calibrated refusal across six categories.
FrontierCode
activePrivate maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.
FrontierMath
activeOriginal, unpublished research-level math problems by expert mathematicians — mostly private to resist contamination. Frontier models solved under 2% at launch.
AgentIF
nearing-saturation707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints — whether instruction following survives deployment-scale prompts.
GPQA
nearing-saturationPhD-written science questions so hard that skilled non-experts with Google score 34% — the 'Google-proof' exam, reported on its 198-question Diamond subset.
GSM-Hard
deprecatedGSM8K test problems with unusually large numerical values, designed to separate reasoning from easy arithmetic.
GSM8K
saturated8.5K grade-school math word problems (2–8 arithmetic steps) — the default elementary-math benchmark and the classic chain-of-thought demonstration.
HellaSwag
saturatedChoose the plausible continuation of an everyday event from one real ending and three adversarially filtered machine generations.
HumanEval
saturatedOpenAI's 164 hand-written Python function-synthesis tasks, graded by executing generated completions against unit tests.
HumanEval+
saturatedHumanEval's 164 Python tasks re-evaluated with roughly 80 times more tests to expose plausible but incorrect programs.
Humanity's Last Exam
active~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models — designed to be the final closed-ended academic benchmark.
IFEval
deprecated~500 prompts, each carrying an explicit, machine-checkable constraint (format, content, or style) that a response must respect.
InfiniteBench
nearing-saturation100k+ token benchmark mixing realistic and synthetic tasks across English, Chinese, code, math, and retrieval.
LiveCodeBench
saturatedContinuously updated contest problems for time-windowed code generation, execution, test prediction, and self-repair evaluation.
LongBench
nearing-saturationBilingual multitask suite for long-context understanding across QA, summarization, few-shot learning, synthetic tasks, and code completion.
LongBench-Pro
activeRealistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.
MATH
saturated12,500 AMC/AIME-level competition problems with worked solutions — usually evaluated on its 500-problem MATH-500 subset, the reasoning-model math standard.
MathArena
activeA continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.
Evaluated AI Foundation Models & LLMs
Claude Fable 5
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Sol
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Fable 5.1 (Hybrid Reasoning)
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 5
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-6 Astra (max)
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 Thinking
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek V4.1 Flash
DeepSeekFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi K3
Moonshot AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Grok 4
xAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.7 Flash
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Grok 4.6
xAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Sonnet 5
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Pro
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.7 Sonnet
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek V4-Pro
DeepSeekFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.6 Flash
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-235B
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.5 Pro
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-R1
DeepSeekFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
QwQ Plus
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-5.2
Zhipu AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 4 Maverick
Meta AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Sonar Reasoning Pro
Perplexity AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Ollama DeepSeek-R1 (Q4_K_M)
Ollama (Inference Runtime)Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Grok 3
xAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MiniMax M3
MiniMaxFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3-mini
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 4 Scout
Meta AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek V4-Flash
DeepSeekFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.0 Flash Thinking
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.5 Flash-Lite
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4.5
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.0 Pro
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
QwQ 32B
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Mistral Large 3
Mistral AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen 2.5 Max
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen 2.5 72B Instruct
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3.3 70B Instruct
Meta AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Hunyuan-Large
TencentFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 3 27B Preview
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Mistral Large 2
Mistral AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi k1.5
Moonshot AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Sonar Pro
Perplexity AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Ollama Llama 3.3 (Q4_K_M)
Ollama (Inference Runtime)Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Yi-Lightning
01.AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Amazon Nova Pro
Amazon AWSFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-V3
DeepSeekFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.0 Flash
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MiniMax-Text-01
MiniMaxFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3.1 405B
Meta AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen 2.5 Coder 32B
Alibaba Cloud / QwenFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ERNIE 4.0 Turbo
BaiduFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Sonar Large
Perplexity AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-4-Plus
Zhipu AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3.3 70B
Meta AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Reka Core
Reka AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Mistral Large 2
Mistral AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codestral 25.01
Mistral AIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 2 27B
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Solar Pro
UpstageFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3 Opus
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 1.5 Pro
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DBRX Instruct
DatabricksFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Haiku
AnthropicFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Snowflake Arctic
SnowflakeFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Phi-4
Microsoft ResearchFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Amazon Nova Lite
Amazon AWSFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Command R+
CohereFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o mini
OpenAIFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 1.5 Flash
Google DeepMindFlagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
QwQ-32B Preview
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.5 Flash (Thinking)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet (v2)
Anthropicflagship tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Sonar
Perplexity AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 2 9B
Google DeepMindsmall tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 2 2B
Google DeepMindsmall tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4.5 (pure LLM)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
r1 / r1-zero (single CoT)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ARChitects (Kaggle 2024 winner)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3-mini high (Epoch independent eval, Tiers 1–3, tools-on)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Goedel-Prover (Lean, 7/644, pass@512)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3-mini (CodeSOTA MBPP pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-V3 (base, 0-shot)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek V3 (open weight)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-V3 (CodeSOTA MBPP pass@1)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3-preview-low (CoT + search/synthesis)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o3 (AIME 2024, pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Jeremy Berman system
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-Coder-32B-Instruct (EvalPlus, greedy, HumanEval base pass@1)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-Coder-32B-Instruct
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-Coder-32B-Instruct (CodeSOTA MBPP pass@1)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-Coder 32B (CodeSOTA MBPP pass@1)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
best frontier model (Tiers 1–3, tools-on)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Grok Beta
xAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-V3 (November 2024)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-V2.5 (November 2024)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet (Oct 2024; CodeSOTA MBPP pass@1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet (upgraded; Anthropic tools scaffold)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
OpenAI o1-mini
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 1.5 Pro 002
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5 72B (base, 0-shot)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1-preview
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (AIME 2024, pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1 (AIME 2024, pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1-preview (September 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1-mini (September 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1 (MATH-500)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
O1 Preview (Sept 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
O1 Mini (Sept 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (Agentless)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (August 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Jamba-1.5-large (README Avg. snapshot)
AI21 Labssnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini-1.5-pro (README Avg. snapshot)
Google DeepMindsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-14B-Instruct-1M (README Avg. snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-235B-A22B (README Avg. snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-14B (README Avg. snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Jamba-1.5-mini (README Avg. snapshot)
AI21 Labssnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-32B (README Avg. snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
EXAONE-4.0-32B (README Avg. snapshot)
LG AI Researchsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen2.5-7B-Instruct-1M (README Avg. snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4-1106-preview (README Avg. snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Fable 5 (LLM Stats snapshot)
Anthropicsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Mythos Preview (LLM Stats snapshot)
Anthropicsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.8 (LLM Stats snapshot)
Anthropicsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Grok 4.5 (LLM Stats snapshot)
xAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Sol (LLM Stats snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.7 (LLM Stats snapshot)
Anthropicsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Terra (LLM Stats snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Sonnet 5 (LLM Stats snapshot)
Anthropicsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Luna (LLM Stats snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-5.2 (LLM Stats snapshot)
Zhipu AIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Muse Spark 1.1 (LLM Stats snapshot)
Meta AIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3.7 Max (LLM Stats snapshot)
Alibaba Cloud / Qwensnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MiniMax M3 (LLM Stats snapshot)
MiniMaxsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.6 Flash (LLM Stats snapshot)
Google DeepMindsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi K2.6 (LLM Stats snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (LLM Stats snapshot)
OpenAIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-5.1 (LLM Stats snapshot)
Zhipu AIsnapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3.1 405B (5-shot)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o Mini (July 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (paper era, Lean, ~1/640)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet (3-shot CoT)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3.5 Sonnet (June 2024)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-Coder-V2-Instruct
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-Coder-V2-Instruct (CodeSOTA MBPP pass@1)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1-mini
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 4.7 (High)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 Pro (High)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Sol (xHigh)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
NVARC
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 4.7 (Max)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (xHigh)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Sol (Max)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Anthropic Opus 4.6 (Max)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.1 Pro (Preview)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (High)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.8 (High)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
SWE-1.6 (FrontierCode 1.1 Main pass rate)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi K2.7 Code (FrontierCode 1.1 Main pass rate)
Moonshot AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
SWE-1.7 (FrontierCode 1.1 Main pass rate)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (FrontierCode 1.1 Main pass rate)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.8 (FrontierCode 1.1 Main pass rate)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3.7 Max (qwen3-7-max; closed)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3.7 Plus (qwen3-7-plus; closed)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-4.7 (glm-4-7; open weight)
Zhipu AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3.6-27B (open weight)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3.6-35B-A3B (open weight)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini-2.5-Pro (paper Overall)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5 (paper Overall)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-235B-A22B-Thinking-2507 (paper Overall)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-Next-80B-A3B-Thinking (paper Overall)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-R1-0528 (paper Overall)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-R1 (paper Overall)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3-30B-A3B-Thinking-2507 (paper Overall)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude-4-Sonnet (paper Overall)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini-2.5-Flash (paper Overall)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MiniMax-M2 (paper Overall)
MiniMaxstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (xhigh, expected performance)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6-Sol (max, expected performance)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o4-mini (CodeSOTA MBPP pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4 (CodeSOTA MBPP pass@1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4.1 (CodeSOTA MBPP pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Sonnet 4 (CodeSOTA MBPP pass@1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Sol (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Terra (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.6 (LLM Stats MRCR v2 8-needle)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 4 31B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.1 Flash-Lite (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.6 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 4 26B-A4B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 4 12B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.6 Luna (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.4 mini (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.4 nano (LLM Stats MRCR v2 8-needle)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.5 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.1 Pro (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Pro (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 4 E4B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Flash (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3.5 Flash-Lite (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 4 E2B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 2.5 Pro Preview 06-05 (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemma 3 27B (LLM Stats MRCR v2 8-needle)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5 (Lean, ~42/660, pass@1, ~10-turn ReAct)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Pro (Challenge Avg@3)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5.5 high (Challenge Avg@3)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Qwen3 32B + SWE-agent (public, launch revision)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o + SWE-agent (public, launch revision)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Sonnet 4 + SWE-agent (public, launch revision)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.1 + SWE-agent (public, launch revision)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5 + SWE-agent (public, launch revision)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 4.5 Opus (high reasoning; mini-SWE-agent 2.0.0)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Flash (high reasoning; mini-SWE-agent 2.0.0)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MiniMax M2.5 (high reasoning; mini-SWE-agent 2.0.0)
MiniMaxstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Opus 4.6 (mini-SWE-agent 2.0.0)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5-2 Codex (mini-SWE-agent 2.0.0)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GLM-5 (high reasoning; mini-SWE-agent 2.0.0)
Zhipu AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5-2 (high reasoning; mini-SWE-agent 2.0.0)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 4.5 Sonnet (high reasoning; mini-SWE-agent 2.0.0)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi K2.5 (high reasoning; mini-SWE-agent 2.0.0)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek V3.2 (high reasoning; mini-SWE-agent 2.0.0)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 3 Pro (mini-SWE-agent 2.0.0)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 4.5 Haiku (high reasoning; mini-SWE-agent 2.0.0)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-5 Mini (mini-SWE-agent 2.0.0)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (198-task Diamond offline)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
o1 (198-task Diamond offline)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Code + Fable 5 (xhigh; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codex + GPT-5.5 (xhigh; TB 2.1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Terminus 2 + Fable 5 (high; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Cursor CLI + Grok 4.5 (high; TB 2.1)
xAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Code + Opus 4.8 (high; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codex + GPT-5.6 Terra (max; TB 2.1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Terminus 2 + GPT-5.5 (xhigh; TB 2.1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
mini-SWE-agent + Muse Spark 1.1 (xhigh; TB 2.1)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codex + GPT-5.6 Luna (max; TB 2.1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Code + Sonnet 5 (high; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Terminus 2 + Gemini 3 Pro (high; TB 2.1)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Code + Opus 4.7 (max; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Terminus 2 + Opus 4.7 (max; TB 2.1)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini CLI + Gemini 3 Pro (high; TB 2.1)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini CLI + Gemini 3.1 Pro (high; TB 2.1)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Terminus 2 + Gemini 3.1 Pro (high; TB 2.1)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude Code + GLM-5.1 (max; TB 2.1)
Zhipu AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (few-shot)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4o (full benchmark)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 Turbo + SWE-agent
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini 1.5 Pro (3-shot CoT)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3 70B (7-shot)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
llama 3 70b-instruct
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
llama 3 8b-instruct
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3 70B
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
CodeQwen1.5-7B-Chat
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
CodeQwen1.5-7B base (Mercury-eval Overall Beyond, 5 samples)
Alibaba Cloud / Qwenstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 3 70B (5-shot)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4-Turbo (April 2024)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4-Turbo
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3 Opus (3-shot CoT)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3 Opus (BM25 retrieval)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 3 Opus (5-shot)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
StarCoder2-15B base (Mercury-eval Overall Beyond, 5 samples)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (paper Table 3 Average)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
YaRN-Mistral (paper Table 3 Average)
Mistral AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Kimi-Chat (paper Table 3 Average)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 2 (paper Table 3 Average)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4V (full benchmark)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
UN codellama-70b-instruct
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini Ultra 1.0 (3-shot CoT)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gemini Ultra (10-shot decontaminated)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PaLM 540B (translate-to-English CoT)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4-Turbo (November 2023)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeepSeek-Coder-33B base (Mercury-eval Overall Beyond, 5 samples)
DeepSeekstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude 2 (BM25 retrieval)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3.5-Turbo-16k (paper Table 3 OverAll)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama2-7B-chat-4k (paper Table 3 OverAll)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
LongChat-v1.5-7B-32k (paper Table 3 OverAll)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
XGen-7B-8k (paper Table 3 OverAll)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
InternLM-7B-8k (paper Table 3 OverAll)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ChatGLM2-6B (paper Table 3 OverAll)
Zhipu AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ChatGLM2-6B-32k (paper Table 3 OverAll)
Zhipu AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Vicuna-v1.5-7B-16k (paper Table 3 OverAll)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
UN codellama-34b-instruct
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
UN codellama-13b-instruct
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
CodeLlama-34B base (Mercury-eval Overall Beyond, 5 samples)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Llama 2 70B (5-shot)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
T0pp (paper Table 3 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Flan-T5 (paper Table 3 Avg)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Flan-UL2 (paper Table 3 Avg)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DaVinci003 (paper Table 3 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ChatGPT (paper Table 3 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Claude (paper Table 3 Avg)
Anthropicstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (paper Table 3 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
CPACE
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (EvalPlus Table 3, greedy, HumanEval base pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (May 2023)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Chameleon (GPT-4, Text-GT, tools)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
text-davinci-003
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ChatGPT (gpt-3.5-turbo)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (3-shot CoT)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (5-shot CoT)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 base (10-shot)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (full MATH)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-4 (5-shot)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
gpt-3.5-turbo
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
LLaMA-65B
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
SantaCoder-1.1B (MultiPL-HumanEval Python, pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PaLM 540B
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codex (code-davinci-002)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Vega v2
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3 + PromptPG (2-shot CoT, Text-GT)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
InCoder-6.7B (MultiPL-HumanEval Python, pass@1)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
code-davinci-002 (MultiPL-HumanEval Python, pass@1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PaLM 540B (5-shot)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Chinchilla 70B
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PaLM 540B (CoT + self-consistency)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ST-MoE-32B
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
AlphaCode 9B (test set, 10@100k, no clustering)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
AlphaCode 41B (test set, 10@100k, no clustering)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
AlphaCode 41B + clustering (test set, 10@100k)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PaLM 540B (8-shot CoT)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Naive (paper Table 2 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BART 256 (paper Table 2 Avg)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BART 512 (paper Table 2 Avg)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BART 1024 (paper Table 2 Avg)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
LED 1024 (paper Table 2 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
LED 4096 (paper Table 2 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
LED 16384 (paper Table 2 Avg)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Gopher 280B
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
KEAR ensemble
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
KEAR single model
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3 175B (fine-tuned)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ERNIE 3.0
Baidustandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Codex-12B (original paper, 164-task HumanEval pass@1 estimator)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Random (paper Table 3 R-1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Ext. Oracle (paper Table 3 R-1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
TextRank (paper Table 3 R-1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PGNet (paper Table 3 R-1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BART (paper Table 3 R-1)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
HMNet* (paper Table 3 R-1)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
PGNet (gold spans) (paper Table 3 R-1)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BART (gold spans) (paper Table 3 R-1)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
HMNet (gold spans) (paper Table 3 R-1)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3 175B (full MATH)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Seq2Seq (CodeSearchNet summarization, six-language overall BLEU)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
Transformer (CodeSearchNet summarization, six-language overall BLEU)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
RoBERTa encoder (CodeSearchNet summarization, six-language overall BLEU)
Meta AIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
CodeBERT encoder (CodeSearchNet summarization, six-language overall BLEU)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeBERTa-xxlarge (1.5B)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeBERTa (TuringNLRv4)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeBERTa ensemble (TuringNLRv4)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
DeBERTa 1.5B (single model)
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3 175B
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
GPT-3 175B (few-shot)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ELECTRA-Large
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
T5 (ensemble)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
T5-11B
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
ALBERT (ensemble)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
RoBERTa-large
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
RoBERTa ensemble
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
RoBERTa
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BERT-Large (fine-tuned)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
RoBERTa-large (fine-tuned)
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
XLNet (ensemble)
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BERT++
OpenAIstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
MT-DNN
Microsoft Researchstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
BERT-Large
Google DeepMindstandard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.
AI Research Labs & Model Creators
01.AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for 01.AI.
AI21 Labs
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for AI21 Labs.
Alibaba Cloud / Qwen
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Alibaba Cloud / Qwen.
Allen Institute for AI (AI2)
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Allen Institute for AI (AI2).
Amazon AWS
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Amazon AWS.
Anthropic
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Anthropic.
Apple
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Apple.
Baidu
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Baidu.
ByteDance Research
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for ByteDance Research.
Center for AI Safety
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Center for AI Safety.
Cohere
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Cohere.
Databricks
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Databricks.
DeepSeek
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for DeepSeek.
Epoch AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Epoch AI.
Google DeepMind
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Google DeepMind.
LG AI Research
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for LG AI Research.
Meta AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Meta AI.
Microsoft Research
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Microsoft Research.
MiniMax
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for MiniMax.
Mistral AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Mistral AI.
Moonshot AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Moonshot AI.
Ollama (Inference Runtime)
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Ollama (Inference Runtime).
OpenAI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for OpenAI.
Perplexity AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Perplexity AI.
Reka AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Reka AI.
Scale AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Scale AI.
Snowflake
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Snowflake.
Stanford CRFM / HELM
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Stanford CRFM / HELM.
Tencent
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Tencent.
Tsinghua University
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Tsinghua University.
UC Berkeley / LMSYS
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for UC Berkeley / LMSYS.
Upstage
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Upstage.
xAI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for xAI.
Zhipu AI
Lab ProfileFrontier AI research, evaluated model portfolio, and benchmark score provenance for Zhipu AI.
Forensic Governance, Feeds & Legal Notices
Contact Research Team
SupportInquiries regarding benchmark methodology audits, dataset submissions, and partnerships.
Submit Feedback & Audit Discrepancies
AuditingReport benchmark contamination anomalies, score discrepancies, or missing models.
Privacy Policy
LegalData collection policies, session telemetry, privacy protections, and user confidentiality.
Terms of Service
LegalPlatform usage terms, API redistribution rules, and community evaluation standards.
Legal & Academic Attribution Notice
AttributionDataset citations, intellectual property protocols, and academic citation standards.
Machine-Readable Manifest (llms.txt)
Machine ReadableCurated plain-text platform index designed specifically for LLM and AI assistant ingestion.
Comprehensive Manifest (llms-full.txt)
Machine ReadableFull forensic evaluation catalog formatted for LLM knowledge retrieval.
XML Sitemap Feed (sitemap.xml)
XML FeedStandard XML sitemap protocol detailing all platform URLs, change frequencies, and priorities.
Crawler Rules (robots.txt)
RobotsSearch engine and AI crawler access rules with explicit permissions for research bots.