MMLU-Redux
A manual error audit of MMLU: experts re-annotated thousands of questions and found ~6.5% are broken — the evidence base for MMLU's label-noise ceiling.
As a competition benchmark, MMLU-Redux is a corrected slice of MMLU: it inherits MMLU's contamination and frontier convergence, so it does not differentiate top models — saturated by CLAUDE.md's definition (usefulness as a model-differentiating eval). Its ENDURING value is different and non-competitive: it remains an actively-used forensic reference and audit methodology for measuring label noise. Status here reflects the eval role; the still-active reference role is documented in prose. saturated_date is set to its launch — it described an already-saturated parent from birth.
Performance Timeline
Longitudinal progression of model scores against human baselines.Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Reference expert baseline adjusted for the 5.7% annotation error rate detected in the original suite.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy on the re-annotated (error-corrected) MMLU questions (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks mmlu-redux --batch_size autoopencompass --datasets mmlu-redux --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for MMLU-Redux, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace