HumanEval
OpenAI's 164 hand-written Python function-synthesis tasks, graded by executing generated completions against unit tests.
Modern coding models routinely cluster near the ceiling on the small public set, while EvalPlus showed the original sparse tests accept materially incorrect programs.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| Qwen2.5-Coder-32B-Instruct (EvalPlus, greedy, HumanEval base pass@1) | 92.7% | 2024-11-12 | vendor-reported | Source ↗ |
| GPT-4 (EvalPlus Table 3, greedy, HumanEval base pass@1) | 88.4% | 2023-05-02 | independent | Source ↗ |
| Codex-12B (original paper, 164-task HumanEval pass@1 estimator) | 28.81% | 2021-07-07 | vendor-reported | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Pass@1 solve rate achieved by human software engineers on single-function Python docstrings without test feedback.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.pass@1 (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks humaneval --batch_size autoopencompass --datasets humaneval --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for HumanEval, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace