BIG-Bench Hard
23 BIG-Bench tasks where models trailed humans — the benchmark that proved chain-of-thought prompting works, then fell to the reasoning-model era.
Frontier and reasoning-tuned models exceed ~90% on the 23-task average and are near-perfect on many individual tasks; the reasoning-model era made BBH trivial. Google DeepMind built BBEH (2025) — replacing each task with a harder counterpart — precisely because BBH stopped discriminating at the top.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| DeepSeek-V3 | 87.5% | 2024-12-26 | vendor-reported | Source ↗ |
| Claude 3.5 Sonnet (3-shot CoT) | 93.1% | 2024-06-21 | vendor-reported | Source ↗ |
| Gemini 1.5 Pro (3-shot CoT) | 89.2% | 2024-05-01 | vendor-reported | Source ↗ |
| Claude 3 Opus (3-shot CoT) | 86.8% | 2024-03-04 | vendor-reported | Source ↗ |
| Gemini Ultra 1.0 (3-shot CoT) | 83.6% | 2023-12-06 | vendor-reported | Source ↗ |
| GPT-4 (3-shot CoT) | 83.1% | 2023-03-14 | vendor-reported | Source ↗ |
| PaLM 540B | 65.2% | 2022-10-17 | independent | Source ↗ |
| Codex (code-davinci-002) | 73.9% | 2022-10-17 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human performance across the 23 hardest algorithmic and multi-step deduction subsets.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.exact-match accuracy, averaged over 23 tasks (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks big-bench-hard --batch_size autoopencompass --datasets big-bench-hard --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for BIG-Bench Hard, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace