MATH
12,500 AMC/AIME-level competition problems with worked solutions — usually evaluated on its 500-problem MATH-500 subset, the reasoning-model math standard.
MATH resisted far longer than GSM8K: GPT-4 managed only ~42% on the full set in 2023. But the reasoning-model era ended it — o1 (2024) reached ~94.8% on MATH-500 and later models sit in the 90s, so it no longer discriminates at the frontier. saturated_date 2024-09 marks the o1 inflection. As with GSM8K, saturated ≠ solved: exact-answer scoring hides whether the reasoning was valid, and headroom has moved to AIME and FrontierMath.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Average human score among elite high school math students across 7 AMC/AIME subdomains.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy on MATH-500 — math-equivalence match on the boxed answer (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks math --batch_size autoopencompass --datasets math --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for MATH, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace