Top Benchmarks & Trusted AI Evaluation Directory

Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.

Capability Domains
Status
Contamination
Sort By
ย 
Displaying 1 of 1 benchmarks (filtered from 63)
SaturatedSafety & Alignment

TruthfulQA

817 adversarial questions probing whether models repeat common human misconceptions โ€” famous for finding that bigger models were often less truthful.

Trust Score
25
Contaminationhigh