Platform Index • Real-Time DynamicUSA & India Dual-Region Optimized

TrustTheBench Platform Sitemap & Directory Index

Comprehensive visual index of verified AI benchmark suites, foundation models, research labs, capability insights, and forensic evaluation datasets. All links are crawlable and refreshed hourly.

63+Audited Benchmarks
372+Evaluated Models
34+AI Research Labs
6Capability Domains
100%Verified Provenance

Core Platform Hubs & Interactive Tools

8 Routes

Capability Domains & Research Insights

6 Articles

AI Benchmark Suites Directory

63 Suites

QuALITY

nearing-saturation

Multiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.

MBPP+

nearing-saturation

A cleaned 378-task MBPP evaluation with roughly 35 times more tests for stricter Python functional correctness.

Mercury

deprecated

Python code-synthesis benchmark that rewards both functional correctness and runtime efficiency against real solution distributions.

MGSM

nearing-saturation

A human-translated 250-problem GSM8K subset for comparing grade-school mathematical reasoning across ten languages.

MMLU

saturated

57-subject multiple-choice exam spanning STEM, humanities, and social sciences — the defining knowledge benchmark of the 2020–2024 era.

MMLU-Pro

nearing-saturation

MMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.

MMLU-Redux

saturated

A manual error audit of MMLU: experts re-annotated thousands of questions and found ~6.5% are broken — the evidence base for MMLU's label-noise ceiling.

MT-Bench

nearing-saturation

80 curated multi-turn questions across 8 skills, each a two-turn conversation scored by a strong LLM judge on a 1–10 scale.

MultiPL-E

deprecated

Parallel HumanEval and MBPP translations for comparing code generation across many programming languages.

Needle-in-a-Haystack

saturated

Synthetic retrieval stress test that hides facts at different depths and context lengths, then asks the model to recover them.

OlympiadBench

active

Bilingual Olympiad and Gaokao mathematics and physics problems, including diagrams, open answers, and proofs.

Omni-MATH

deprecated

A 4,428-problem Olympiad mathematics suite spanning more than 33 subdomains and ten annotated difficulty levels.

OpenAI-MRCR

active

MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.

PutnamBench

active

Formal theorem-proving on Putnam problems — models write a complete Lean/Isabelle/Coq proof, machine-checked by the proof assistant's kernel.

QASPER

nearing-saturation

Scientific-paper QA over full NLP papers, with abstractive, extractive, yes/no, and unanswerable answers plus supporting evidence.

QMSum

nearing-saturation

Query-focused meeting summarization over long transcripts from academic, product, and committee meetings.

MBPP

saturated

Mostly Basic Programming Problems: 974 crowd-sourced Python synthesis tasks intended for entry-level programmers.

RULER

nearing-saturation

Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.

SCROLLS

nearing-saturation

Seven-task long-sequence suite covering summarization, QA, and NLI over naturally long English texts.

SimpleQA

active

4,326 short fact-seeking questions chosen to stump GPT-4 — a hallucination test graded correct / incorrect / not-attempted.

Soohak

active

Fresh research-level mathematics authored by mathematicians, plus unanswerable prompts that reward justified refusal.

StructFlowBench

nearing-saturation

155 multi-turn dialogues scored on fine-grained constraints plus six inter-turn dialogue relationships (recall, refinement, expansion, follow-up, summary).

SuperGLUE

saturated

Eight English NLU tasks spanning question answering, inference, causal reasoning, word sense, and coreference.

SWE-bench

deprecated

Real GitHub issues from 12 Python repositories, solved by generating patches that must pass repository tests.

SWE-bench Pro

active

Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.

SWE-bench Verified

saturated

A 500-task, engineer-reviewed subset of SWE-bench that became the standard repository-level coding-agent test.

SWE-Lancer

deprecated

Current 198-task offline patch benchmark derived from paid Upwork work, with historical mixed implementation and manager-task variants.

TabMWP

saturated

Grade-level math word problems that require combining a question with structured, textual, or image-rendered tables.

Terminal-Bench

active

Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.

TruthfulQA

saturated

817 adversarial questions probing whether models repeat common human misconceptions — famous for finding that bigger models were often less truthful.

WinoGrande

saturated

Binary fill-in-the-blank commonsense problems, scaled from Winograd schemas and adversarially filtered to reduce dataset shortcuts.

ZeroSCROLLS

nearing-saturation

Zero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.

GLUE

saturated

Nine-task English NLU suite combining acceptability, sentiment, similarity, paraphrase, inference and coreference into one score.

AGIEval

saturated

Standardized human exams — SATs, LSATs, China's Gaokao and civil-service tests — repurposed to grade models against real human test-takers.

AIME

nearing-saturation

The MAA's competition exam — 15 integer-answer problems per test, refreshed yearly so each new exam is briefly contamination-free.

ARC-AGI

saturated

Colored-grid abstraction tasks that test few-shot rule induction, not memorized facts, language fluency, or school math.

ARC-AGI-2

nearing-saturation

The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.

ARC-AGI-3

active

Interactive ARC environments that score agents by human-relative action efficiency, planning, exploration, and adaptation.

BIG-Bench

deprecated

204-task community-built mega-suite spanning reasoning, knowledge, math, code and social bias — the sprawling proving ground BBH was distilled from.

BIG-Bench Hard

saturated

23 BIG-Bench tasks where models trailed humans — the benchmark that proved chain-of-thought prompting works, then fell to the reasoning-model era.

CodeContests

deprecated

DeepMind's competitive-programming corpus with temporally split problems, human submissions, and generated correctness tests.

CodeXGLUE

deprecated

A broad code-intelligence suite spanning understanding, retrieval, completion, translation, repair, generation, and summarization.

CommonsenseQA

saturated

Five-choice questions derived from ConceptNet relations, designed to require everyday knowledge beyond a supplied passage.

ComplexBench

nearing-saturation

Bilingual benchmark of how models follow instructions when several constraints must hold at once — where stacking constraints is not linear.

CyberSecEval

deprecated

Meta's evolving suite for insecure code, cyberattack assistance, prompt injection, exploitation, SOC analysis, and autopatching.

FireBench

active

Evaluates instruction following for enterprise/API pipelines — exact output format, strict step ordering, ranking, and calibrated refusal across six categories.

FrontierCode

active

Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.

FrontierMath

active

Original, unpublished research-level math problems by expert mathematicians — mostly private to resist contamination. Frontier models solved under 2% at launch.

AgentIF

nearing-saturation

707 instructions from 50 real-world agents, averaging 1,723 words and 11.9 constraints — whether instruction following survives deployment-scale prompts.

GPQA

nearing-saturation

PhD-written science questions so hard that skilled non-experts with Google score 34% — the 'Google-proof' exam, reported on its 198-question Diamond subset.

GSM-Hard

deprecated

GSM8K test problems with unusually large numerical values, designed to separate reasoning from easy arithmetic.

GSM8K

saturated

8.5K grade-school math word problems (2–8 arithmetic steps) — the default elementary-math benchmark and the classic chain-of-thought demonstration.

HellaSwag

saturated

Choose the plausible continuation of an everyday event from one real ending and three adversarially filtered machine generations.

HumanEval

saturated

OpenAI's 164 hand-written Python function-synthesis tasks, graded by executing generated completions against unit tests.

HumanEval+

saturated

HumanEval's 164 Python tasks re-evaluated with roughly 80 times more tests to expose plausible but incorrect programs.

Humanity's Last Exam

active

~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models — designed to be the final closed-ended academic benchmark.

IFEval

deprecated

~500 prompts, each carrying an explicit, machine-checkable constraint (format, content, or style) that a response must respect.

InfiniteBench

nearing-saturation

100k+ token benchmark mixing realistic and synthetic tasks across English, Chinese, code, math, and retrieval.

LiveCodeBench

saturated

Continuously updated contest problems for time-windowed code generation, execution, test prediction, and self-repair evaluation.

LongBench

nearing-saturation

Bilingual multitask suite for long-context understanding across QA, summarization, few-shot learning, synthetic tasks, and code completion.

LongBench-Pro

active

Realistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.

MATH

saturated

12,500 AMC/AIME-level competition problems with worked solutions — usually evaluated on its 500-problem MATH-500 subset, the reasoning-model math standard.

MathArena

active

A continuously refreshed platform evaluating models on newly released math competitions, proofs, research problems, and formalization.

Evaluated AI Foundation Models & LLMs

372 Models

Claude Fable 5

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Sol

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Fable 5.1 (Hybrid Reasoning)

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 5

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-6 Astra (max)

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 Thinking

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek V4.1 Flash

DeepSeek

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi K3

Moonshot AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Grok 4

xAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.7 Flash

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Grok 4.6

xAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Sonnet 5

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Pro

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.7 Sonnet

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek V4-Pro

DeepSeek

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.6 Flash

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-235B

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.5 Pro

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-R1

DeepSeek

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

QwQ Plus

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-5.2

Zhipu AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 4 Maverick

Meta AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Sonar Reasoning Pro

Perplexity AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Ollama DeepSeek-R1 (Q4_K_M)

Ollama (Inference Runtime)

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Grok 3

xAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MiniMax M3

MiniMax

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3-mini

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 4 Scout

Meta AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek V4-Flash

DeepSeek

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.0 Flash Thinking

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.5 Flash-Lite

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4.5

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.0 Pro

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

QwQ 32B

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Mistral Large 3

Mistral AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen 2.5 Max

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen 2.5 72B Instruct

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3.3 70B Instruct

Meta AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Hunyuan-Large

Tencent

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 3 27B Preview

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Mistral Large 2

Mistral AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi k1.5

Moonshot AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Sonar Pro

Perplexity AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Ollama Llama 3.3 (Q4_K_M)

Ollama (Inference Runtime)

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Yi-Lightning

01.AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Amazon Nova Pro

Amazon AWS

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-V3

DeepSeek

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.0 Flash

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MiniMax-Text-01

MiniMax

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3.1 405B

Meta AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen 2.5 Coder 32B

Alibaba Cloud / Qwen

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ERNIE 4.0 Turbo

Baidu

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Sonar Large

Perplexity AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-4-Plus

Zhipu AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3.3 70B

Meta AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Reka Core

Reka AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Mistral Large 2

Mistral AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codestral 25.01

Mistral AI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 2 27B

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Solar Pro

Upstage

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3 Opus

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 1.5 Pro

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DBRX Instruct

Databricks

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Haiku

Anthropic

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Snowflake Arctic

Snowflake

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Phi-4

Microsoft Research

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Amazon Nova Lite

Amazon AWS

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Command R+

Cohere

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o mini

OpenAI

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 1.5 Flash

Google DeepMind

Flagship frontier model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

QwQ-32B Preview

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.5 Flash (Thinking)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet (v2)

Anthropic

flagship tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Sonar

Perplexity AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 2 9B

Google DeepMind

small tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 2 2B

Google DeepMind

small tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4.5 (pure LLM)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

r1 / r1-zero (single CoT)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ARChitects (Kaggle 2024 winner)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3-mini high (Epoch independent eval, Tiers 1–3, tools-on)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Goedel-Prover (Lean, 7/644, pass@512)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3-mini (CodeSOTA MBPP pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-V3 (base, 0-shot)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek V3 (open weight)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-V3 (CodeSOTA MBPP pass@1)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3-preview-low (CoT + search/synthesis)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o3 (AIME 2024, pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Jeremy Berman system

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-Coder-32B-Instruct (EvalPlus, greedy, HumanEval base pass@1)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-Coder-32B-Instruct

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-Coder-32B-Instruct (CodeSOTA MBPP pass@1)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-Coder 32B (CodeSOTA MBPP pass@1)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

best frontier model (Tiers 1–3, tools-on)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Grok Beta

xAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-V3 (November 2024)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-V2.5 (November 2024)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet (Oct 2024; CodeSOTA MBPP pass@1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet (upgraded; Anthropic tools scaffold)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

OpenAI o1-mini

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 1.5 Pro 002

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5 72B (base, 0-shot)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1-preview

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (AIME 2024, pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1 (AIME 2024, pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1-preview (September 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1-mini (September 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1 (MATH-500)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

O1 Preview (Sept 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

O1 Mini (Sept 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (Agentless)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (August 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Jamba-1.5-large (README Avg. snapshot)

AI21 Labs

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini-1.5-pro (README Avg. snapshot)

Google DeepMind

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-14B-Instruct-1M (README Avg. snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-235B-A22B (README Avg. snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-14B (README Avg. snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Jamba-1.5-mini (README Avg. snapshot)

AI21 Labs

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-32B (README Avg. snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

EXAONE-4.0-32B (README Avg. snapshot)

LG AI Research

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen2.5-7B-Instruct-1M (README Avg. snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4-1106-preview (README Avg. snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Fable 5 (LLM Stats snapshot)

Anthropic

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Mythos Preview (LLM Stats snapshot)

Anthropic

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.8 (LLM Stats snapshot)

Anthropic

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Grok 4.5 (LLM Stats snapshot)

xAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Sol (LLM Stats snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.7 (LLM Stats snapshot)

Anthropic

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Terra (LLM Stats snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Sonnet 5 (LLM Stats snapshot)

Anthropic

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Luna (LLM Stats snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-5.2 (LLM Stats snapshot)

Zhipu AI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Muse Spark 1.1 (LLM Stats snapshot)

Meta AI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3.7 Max (LLM Stats snapshot)

Alibaba Cloud / Qwen

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MiniMax M3 (LLM Stats snapshot)

MiniMax

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.6 Flash (LLM Stats snapshot)

Google DeepMind

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi K2.6 (LLM Stats snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (LLM Stats snapshot)

OpenAI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-5.1 (LLM Stats snapshot)

Zhipu AI

snapshot tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3.1 405B (5-shot)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o Mini (July 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (paper era, Lean, ~1/640)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet (3-shot CoT)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3.5 Sonnet (June 2024)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-Coder-V2-Instruct

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-Coder-V2-Instruct (CodeSOTA MBPP pass@1)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1-mini

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 4.7 (High)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 Pro (High)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Sol (xHigh)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

NVARC

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 4.7 (Max)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (xHigh)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Sol (Max)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Anthropic Opus 4.6 (Max)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.1 Pro (Preview)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (High)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.8 (High)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

SWE-1.6 (FrontierCode 1.1 Main pass rate)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi K2.7 Code (FrontierCode 1.1 Main pass rate)

Moonshot AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

SWE-1.7 (FrontierCode 1.1 Main pass rate)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (FrontierCode 1.1 Main pass rate)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.8 (FrontierCode 1.1 Main pass rate)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3.7 Max (qwen3-7-max; closed)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3.7 Plus (qwen3-7-plus; closed)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-4.7 (glm-4-7; open weight)

Zhipu AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3.6-27B (open weight)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3.6-35B-A3B (open weight)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini-2.5-Pro (paper Overall)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5 (paper Overall)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-235B-A22B-Thinking-2507 (paper Overall)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-Next-80B-A3B-Thinking (paper Overall)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-R1-0528 (paper Overall)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-R1 (paper Overall)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3-30B-A3B-Thinking-2507 (paper Overall)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude-4-Sonnet (paper Overall)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini-2.5-Flash (paper Overall)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MiniMax-M2 (paper Overall)

MiniMax

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (xhigh, expected performance)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6-Sol (max, expected performance)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o4-mini (CodeSOTA MBPP pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4 (CodeSOTA MBPP pass@1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4.1 (CodeSOTA MBPP pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Sonnet 4 (CodeSOTA MBPP pass@1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Sol (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Terra (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.6 (LLM Stats MRCR v2 8-needle)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 4 31B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.1 Flash-Lite (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.6 Flash (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 4 26B-A4B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 4 12B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.6 Luna (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.4 mini (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.4 nano (LLM Stats MRCR v2 8-needle)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.5 Flash (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.1 Pro (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Pro (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 4 E4B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Flash (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3.5 Flash-Lite (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 4 E2B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 2.5 Pro Preview 06-05 (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemma 3 27B (LLM Stats MRCR v2 8-needle)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5 (Lean, ~42/660, pass@1, ~10-turn ReAct)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Pro (Challenge Avg@3)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5.5 high (Challenge Avg@3)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Qwen3 32B + SWE-agent (public, launch revision)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o + SWE-agent (public, launch revision)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Sonnet 4 + SWE-agent (public, launch revision)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.1 + SWE-agent (public, launch revision)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5 + SWE-agent (public, launch revision)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 4.5 Opus (high reasoning; mini-SWE-agent 2.0.0)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Flash (high reasoning; mini-SWE-agent 2.0.0)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MiniMax M2.5 (high reasoning; mini-SWE-agent 2.0.0)

MiniMax

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Opus 4.6 (mini-SWE-agent 2.0.0)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5-2 Codex (mini-SWE-agent 2.0.0)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GLM-5 (high reasoning; mini-SWE-agent 2.0.0)

Zhipu AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5-2 (high reasoning; mini-SWE-agent 2.0.0)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 4.5 Sonnet (high reasoning; mini-SWE-agent 2.0.0)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi K2.5 (high reasoning; mini-SWE-agent 2.0.0)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek V3.2 (high reasoning; mini-SWE-agent 2.0.0)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 3 Pro (mini-SWE-agent 2.0.0)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 4.5 Haiku (high reasoning; mini-SWE-agent 2.0.0)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-5 Mini (mini-SWE-agent 2.0.0)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (198-task Diamond offline)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

o1 (198-task Diamond offline)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Code + Fable 5 (xhigh; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codex + GPT-5.5 (xhigh; TB 2.1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Terminus 2 + Fable 5 (high; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Cursor CLI + Grok 4.5 (high; TB 2.1)

xAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Code + Opus 4.8 (high; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codex + GPT-5.6 Terra (max; TB 2.1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Terminus 2 + GPT-5.5 (xhigh; TB 2.1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

mini-SWE-agent + Muse Spark 1.1 (xhigh; TB 2.1)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codex + GPT-5.6 Luna (max; TB 2.1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Code + Sonnet 5 (high; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Terminus 2 + Gemini 3 Pro (high; TB 2.1)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Code + Opus 4.7 (max; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Terminus 2 + Opus 4.7 (max; TB 2.1)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini CLI + Gemini 3 Pro (high; TB 2.1)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini CLI + Gemini 3.1 Pro (high; TB 2.1)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Terminus 2 + Gemini 3.1 Pro (high; TB 2.1)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude Code + GLM-5.1 (max; TB 2.1)

Zhipu AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (few-shot)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4o (full benchmark)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 Turbo + SWE-agent

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini 1.5 Pro (3-shot CoT)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3 70B (7-shot)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

llama 3 70b-instruct

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

llama 3 8b-instruct

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3 70B

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

CodeQwen1.5-7B-Chat

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

CodeQwen1.5-7B base (Mercury-eval Overall Beyond, 5 samples)

Alibaba Cloud / Qwen

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 3 70B (5-shot)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4-Turbo (April 2024)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4-Turbo

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3 Opus (3-shot CoT)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3 Opus (BM25 retrieval)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 3 Opus (5-shot)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

StarCoder2-15B base (Mercury-eval Overall Beyond, 5 samples)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (paper Table 3 Average)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

YaRN-Mistral (paper Table 3 Average)

Mistral AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Kimi-Chat (paper Table 3 Average)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 2 (paper Table 3 Average)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4V (full benchmark)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

UN codellama-70b-instruct

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini Ultra 1.0 (3-shot CoT)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gemini Ultra (10-shot decontaminated)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PaLM 540B (translate-to-English CoT)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4-Turbo (November 2023)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeepSeek-Coder-33B base (Mercury-eval Overall Beyond, 5 samples)

DeepSeek

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude 2 (BM25 retrieval)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3.5-Turbo-16k (paper Table 3 OverAll)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama2-7B-chat-4k (paper Table 3 OverAll)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

LongChat-v1.5-7B-32k (paper Table 3 OverAll)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

XGen-7B-8k (paper Table 3 OverAll)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

InternLM-7B-8k (paper Table 3 OverAll)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ChatGLM2-6B (paper Table 3 OverAll)

Zhipu AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ChatGLM2-6B-32k (paper Table 3 OverAll)

Zhipu AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Vicuna-v1.5-7B-16k (paper Table 3 OverAll)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

UN codellama-34b-instruct

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

UN codellama-13b-instruct

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

CodeLlama-34B base (Mercury-eval Overall Beyond, 5 samples)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Llama 2 70B (5-shot)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

T0pp (paper Table 3 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Flan-T5 (paper Table 3 Avg)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Flan-UL2 (paper Table 3 Avg)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DaVinci003 (paper Table 3 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ChatGPT (paper Table 3 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Claude (paper Table 3 Avg)

Anthropic

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (paper Table 3 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

CPACE

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (EvalPlus Table 3, greedy, HumanEval base pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (May 2023)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Chameleon (GPT-4, Text-GT, tools)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

text-davinci-003

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ChatGPT (gpt-3.5-turbo)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (3-shot CoT)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (5-shot CoT)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 base (10-shot)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (full MATH)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-4 (5-shot)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

gpt-3.5-turbo

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

LLaMA-65B

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

SantaCoder-1.1B (MultiPL-HumanEval Python, pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PaLM 540B

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codex (code-davinci-002)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Vega v2

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3 + PromptPG (2-shot CoT, Text-GT)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

InCoder-6.7B (MultiPL-HumanEval Python, pass@1)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

code-davinci-002 (MultiPL-HumanEval Python, pass@1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PaLM 540B (5-shot)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Chinchilla 70B

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PaLM 540B (CoT + self-consistency)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ST-MoE-32B

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

AlphaCode 9B (test set, 10@100k, no clustering)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

AlphaCode 41B (test set, 10@100k, no clustering)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

AlphaCode 41B + clustering (test set, 10@100k)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PaLM 540B (8-shot CoT)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Naive (paper Table 2 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BART 256 (paper Table 2 Avg)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BART 512 (paper Table 2 Avg)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BART 1024 (paper Table 2 Avg)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

LED 1024 (paper Table 2 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

LED 4096 (paper Table 2 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

LED 16384 (paper Table 2 Avg)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Gopher 280B

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

KEAR ensemble

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

KEAR single model

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3 175B (fine-tuned)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ERNIE 3.0

Baidu

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Codex-12B (original paper, 164-task HumanEval pass@1 estimator)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Random (paper Table 3 R-1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Ext. Oracle (paper Table 3 R-1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

TextRank (paper Table 3 R-1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PGNet (paper Table 3 R-1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BART (paper Table 3 R-1)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

HMNet* (paper Table 3 R-1)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

PGNet (gold spans) (paper Table 3 R-1)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BART (gold spans) (paper Table 3 R-1)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

HMNet (gold spans) (paper Table 3 R-1)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3 175B (full MATH)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Seq2Seq (CodeSearchNet summarization, six-language overall BLEU)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

Transformer (CodeSearchNet summarization, six-language overall BLEU)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

RoBERTa encoder (CodeSearchNet summarization, six-language overall BLEU)

Meta AI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

CodeBERT encoder (CodeSearchNet summarization, six-language overall BLEU)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeBERTa-xxlarge (1.5B)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeBERTa (TuringNLRv4)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeBERTa ensemble (TuringNLRv4)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

DeBERTa 1.5B (single model)

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3 175B

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

GPT-3 175B (few-shot)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ELECTRA-Large

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

T5 (ensemble)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

T5-11B

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

ALBERT (ensemble)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

RoBERTa-large

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

RoBERTa ensemble

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

RoBERTa

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BERT-Large (fine-tuned)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

RoBERTa-large (fine-tuned)

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

XLNet (ensemble)

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BERT++

OpenAI

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

MT-DNN

Microsoft Research

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

BERT-Large

Google DeepMind

standard tier foundation model evaluated across LiveBench, SWE-bench, and reasoning benchmarks.

AI Research Labs & Model Creators

34 Labs

01.AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for 01.AI.

AI21 Labs

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for AI21 Labs.

Alibaba Cloud / Qwen

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Alibaba Cloud / Qwen.

Allen Institute for AI (AI2)

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Allen Institute for AI (AI2).

Amazon AWS

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Amazon AWS.

Anthropic

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Anthropic.

Apple

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Apple.

Baidu

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Baidu.

ByteDance Research

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for ByteDance Research.

Center for AI Safety

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Center for AI Safety.

Cohere

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Cohere.

Databricks

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Databricks.

DeepSeek

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for DeepSeek.

Epoch AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Epoch AI.

Google DeepMind

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Google DeepMind.

LG AI Research

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for LG AI Research.

Meta AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Meta AI.

Microsoft Research

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Microsoft Research.

MiniMax

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for MiniMax.

Mistral AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Mistral AI.

Moonshot AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Moonshot AI.

Ollama (Inference Runtime)

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Ollama (Inference Runtime).

OpenAI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for OpenAI.

Perplexity AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Perplexity AI.

Reka AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Reka AI.

Scale AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Scale AI.

Snowflake

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Snowflake.

Stanford CRFM / HELM

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Stanford CRFM / HELM.

Tencent

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Tencent.

Tsinghua University

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Tsinghua University.

UC Berkeley / LMSYS

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for UC Berkeley / LMSYS.

Upstage

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Upstage.

xAI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for xAI.

Zhipu AI

Lab Profile

Frontier AI research, evaluated model portfolio, and benchmark score provenance for Zhipu AI.

Forensic Governance, Feeds & Legal Notices

9 Protocols