SuperGLUE
Eight English NLU tasks spanning question answering, inference, causal reasoning, word sense, and coreference.
The official aggregate human estimate (89.8) was crossed by a single DeBERTa model (89.9) by 2020-12-29; the 2022 leaderboard leader reached 91.3, leaving a narrow, noisy ceiling.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| Vega v2 | 91.3% | 2022-10-08 | vendor-reported | Source ↗ |
| PaLM 540B | 90.4% | 2022-04-04 | vendor-reported | Source ↗ |
| ST-MoE-32B | 91.2% | 2022-02-17 | vendor-reported | Source ↗ |
| ERNIE 3.0 | 90.6% | 2021-07-12 | vendor-reported | Source ↗ |
| DeBERTa ensemble (TuringNLRv4) | 90.3% | 2021-01-06 | vendor-reported | Source ↗ |
| DeBERTa 1.5B (single model) | 89.9% | 2020-12-29 | vendor-reported | Source ↗ |
| T5-11B | 89.3% | 2019-10-23 | vendor-reported | Source ↗ |
| BERT++ | 71.5% | 2019-05-02 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human crowdworker accuracy on SuperGLUE linguistic reasoning and entailment suite.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.SuperGLUE aggregate score (%)Dataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks superglue --batch_size autoopencompass --datasets superglue --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for SuperGLUE, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace