GSM8K
8.5K grade-school math word problems (2–8 arithmetic steps) — the default elementary-math benchmark and the classic chain-of-thought demonstration.
Frontier models cleared ~95% by 2023–24 and the reasoning-model era pushed scores to ~97–99%, collapsing headroom. But saturation here is the sharpest 'saturated ≠ solved' case on the wiki: GSM-Symbolic (Apple, 2024) showed that changing only names/numbers, and especially adding one irrelevant clause (GSM-NoOp), drops accuracy heavily across models — so the SCORE is saturated while the CAPABILITY (robust arithmetic reasoning) is not. saturated_date 2024-01 marks the point frontier scores stopped discriminating; the discriminating signal moved to harder math (MATH → AIME → FrontierMath) and to robustness variants.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.OpenAI measured human problem solvers scoring 96% due to occasional arithmetic and reading slip-ups (unverified crowdworkers score ~80%).
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy — exact match on the final numeric answer (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks gsm8k --batch_size autoopencompass --datasets gsm8k --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for GSM8K, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace