AGIEval
Standardized human exams — SATs, LSATs, China's Gaokao and civil-service tests — repurposed to grade models against real human test-takers.
GPT-4 already cleared the average-human line (67%) on several sections at launch; by 2024 frontier models pushed the 20-task aggregate past the average human and near the top-percentile 91% ceiling, and the human-exam framing lost its discriminative power at the top. saturated_date (2024-06) marks the GPT-4-class era where the average-human baseline stopped separating strong models.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| DeepSeek-V3 (base, 0-shot) | 79.6% | 2024-12-26 | vendor-reported | Source ↗ |
| Qwen2.5 72B (base, 0-shot) | 75.8% | 2024-09-19 | vendor-reported | Source ↗ |
| GPT-4o (few-shot) | 69% | 2024-05-13 | vendor-reported | Source ↗ |
| text-davinci-003 | 37.4% | 2023-04-13 | independent | Source ↗ |
| ChatGPT (gpt-3.5-turbo) | 43.2% | 2023-04-13 | independent | Source ↗ |
| GPT-4 | 58.4% | 2023-04-13 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Average score of human test-takers across SAT, LSAT, GRE, Gaokao, and Chinese civil service exams.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy, macro-averaged over the 20 tasks (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks agieval --batch_size autoopencompass --datasets agieval --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for AGIEval, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace