FrontierCode
Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
Private maintainer-authored repository tasks graded for mergeability: correctness, tests, scope discipline, style, and code quality.
Long-horizon repository tasks across public, held-out, and commercial codebases, with a continuing public model leaderboard.
Eighty-nine realistic terminal tasks graded in containers; version 2.1 repairs 2.0's dependency, resource, and specification defects.
A cleaned 378-task MBPP evaluation with roughly 35 times more tests for stricter Python functional correctness.
Continuously updated contest problems for time-windowed code generation, execution, test prediction, and self-repair evaluation.
Mostly Basic Programming Problems: 974 crowd-sourced Python synthesis tasks intended for entry-level programmers.
A 500-task, engineer-reviewed subset of SWE-bench that became the standard repository-level coding-agent test.
OpenAI's 164 hand-written Python function-synthesis tasks, graded by executing generated completions against unit tests.
HumanEval's 164 Python tasks re-evaluated with roughly 80 times more tests to expose plausible but incorrect programs.
Current 198-task offline patch benchmark derived from paid Upwork work, with historical mixed implementation and manager-task variants.
A broad code-intelligence suite spanning understanding, retrieval, completion, translation, repair, generation, and summarization.
Python code-synthesis benchmark that rewards both functional correctness and runtime efficiency against real solution distributions.
Parallel HumanEval and MBPP translations for comparing code generation across many programming languages.
Real GitHub issues from 12 Python repositories, solved by generating patches that must pass repository tests.
DeepMind's competitive-programming corpus with temporally split problems, human submissions, and generated correctness tests.
Meta's evolving suite for insecure code, cyberattack assistance, prompt injection, exploitation, SOC analysis, and autopatching.