SWE-bench Pro
Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.
The benchmark still has a maintained public leaderboard and substantial score spread: the 2026-08-03 LLM Stats snapshot ranges from 0.584 to 0.800 across these 17 leading models, despite documented task-quality concerns.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| GPT-5.6 Terra (LLM Stats snapshot) | 63.4% | 2024-08-01 | vendor-reported | Source ↗ |
| GLM-5.1 (LLM Stats snapshot) | 58.4% | 2024-08-01 | vendor-reported | Source ↗ |
| GPT-5.5 (LLM Stats snapshot) | 58.6% | 2024-08-01 | vendor-reported | Source ↗ |
| Kimi K2.6 (LLM Stats snapshot) | 58.6% | 2024-08-01 | vendor-reported | Source ↗ |
| Gemini 3.6 Flash (LLM Stats snapshot) | 58.7% | 2024-08-01 | vendor-reported | Source ↗ |
| MiniMax M3 (LLM Stats snapshot) | 59% | 2024-08-01 | vendor-reported | Source ↗ |
| Qwen3.7 Max (LLM Stats snapshot) | 60.6% | 2024-08-01 | vendor-reported | Source ↗ |
| Muse Spark 1.1 (LLM Stats snapshot) | 61.5% | 2024-08-01 | vendor-reported | Source ↗ |
| GLM-5.2 (LLM Stats snapshot) | 62.1% | 2024-08-01 | vendor-reported | Source ↗ |
| GPT-5.6 Luna (LLM Stats snapshot) | 62.7% | 2024-08-01 | vendor-reported | Source ↗ |
| Claude Sonnet 5 (LLM Stats snapshot) | 63.2% | 2024-08-01 | vendor-reported | Source ↗ |
| Claude Opus 4.7 (LLM Stats snapshot) | 64.3% | 2024-08-01 | vendor-reported | Source ↗ |
| GPT-5.6 Sol (LLM Stats snapshot) | 64.6% | 2024-08-01 | vendor-reported | Source ↗ |
| Grok 4.5 (LLM Stats snapshot) | 64.7% | 2024-08-01 | vendor-reported | Source ↗ |
| Claude Opus 4.8 (LLM Stats snapshot) | 69.2% | 2024-08-01 | vendor-reported | Source ↗ |
| Claude Mythos Preview (LLM Stats snapshot) | 77.8% | 2024-08-01 | vendor-reported | Source ↗ |
| Claude Fable 5 (LLM Stats snapshot) | 80% | 2024-08-01 | vendor-reported | Source ↗ |
| Qwen3 32B + SWE-agent (public, launch revision) | 3.4% | 2024-06-01 | independent | Source ↗ |
| GPT-5 + SWE-agent (public, launch revision) | 23.3% | 2024-06-01 | independent | Source ↗ |
| Claude Opus 4.1 + SWE-agent (public, launch revision) | 22.7% | 2024-06-01 | independent | Source ↗ |
| Claude Sonnet 4 + SWE-agent (public, launch revision) | 17.6% | 2024-06-01 | independent | Source ↗ |
| GPT-4o + SWE-agent (public, launch revision) | 4.9% | 2024-06-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Professional software engineers completing repository-scale enterprise feature requests and bug fixes in 4 to 12 hours.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.resolved instances on the public set (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks swe-bench-pro --batch_size autoopencompass --datasets swe-bench-pro --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for SWE-bench Pro, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace