Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 6 of 6 benchmarks (filtered from 63)
ActiveKnowledge

SimpleQA

4,326 short fact-seeking questions chosen to stump GPT-4 โ€” a hallucination test graded correct / incorrect / not-attempted.

Trust Score
90
Contaminationmedium
ActiveKnowledge

Humanity's Last Exam

~2,500 expert-written questions across 100+ subjects, each filtered to stump frontier models โ€” designed to be the final closed-ended academic benchmark.

Trust Score
90
Contaminationmedium
Nearing SaturationKnowledge

MMLU-Pro

MMLU rebuilt to fix its flaws: 10 options instead of 4, trivia and label noise filtered out, and reasoning-heavy questions that reward chain-of-thought.

Trust Score
75
Contaminationmedium
Nearing SaturationKnowledge

GPQA

PhD-written science questions so hard that skilled non-experts with Google score 34% โ€” the 'Google-proof' exam, reported on its 198-question Diamond subset.

Trust Score
75
Contaminationmedium
SaturatedKnowledge

MMLU

57-subject multiple-choice exam spanning STEM, humanities, and social sciences โ€” the defining knowledge benchmark of the 2020โ€“2024 era.

Trust Score
25
Contaminationhigh
SaturatedKnowledge

MMLU-Redux

A manual error audit of MMLU: experts re-annotated thousands of questions and found ~6.5% are broken โ€” the evidence base for MMLU's label-noise ceiling.

Trust Score
25
Contaminationhigh