ARC-AGI-2
The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
The harder static ARC successor: human-calibrated colored-grid abstraction tasks, now pressured by 2026 frontier systems.
Eight English NLU tasks spanning question answering, inference, causal reasoning, word sense, and coreference.
Binary fill-in-the-blank commonsense problems, scaled from Winograd schemas and adversarially filtered to reduce dataset shortcuts.
Nine-task English NLU suite combining acceptability, sentiment, similarity, paraphrase, inference and coreference into one score.
Standardized human exams โ SATs, LSATs, China's Gaokao and civil-service tests โ repurposed to grade models against real human test-takers.
Colored-grid abstraction tasks that test few-shot rule induction, not memorized facts, language fluency, or school math.
23 BIG-Bench tasks where models trailed humans โ the benchmark that proved chain-of-thought prompting works, then fell to the reasoning-model era.
Five-choice questions derived from ConceptNet relations, designed to require everyday knowledge beyond a supplied passage.
Choose the plausible continuation of an everyday event from one real ending and three adversarially filtered machine generations.
204-task community-built mega-suite spanning reasoning, knowledge, math, code and social bias โ the sprawling proving ground BBH was distilled from.