Sufficiently difficult to discriminate frontier reasoning models with test-time compute.
Expected Score50% – 95%
ContaminationLOW
Human Baseline45%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks frontiercode --batch_size auto Benchmark Specs Sufficiently difficult to discriminate frontier reasoning models with test-time compute.
Expected Score50% – 95%
ContaminationMEDIUM
Human Baseline65%
Refresh Cadencecontinuous
lm_eval --model hf --model_args pretrained=<your_model> --tasks --scenario codegeneration --batch_size auto Benchmark Specs Sufficiently difficult to discriminate frontier reasoning models with test-time compute.
Expected Score50% – 95%
ContaminationMEDIUM
Human Baseline74%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks swe_bench_pro_eval.py --batch_size auto Benchmark Specs Sufficiently difficult to discriminate frontier reasoning models with test-time compute.
Expected Score50% – 95%
ContaminationHIGH
Human Baseline78.4%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks SWE-bench_Verified --batch_size auto Benchmark Specs Active benchmark with solid headroom against top foundation models.
Expected Score70% – 90%
ContaminationHIGH
Human Baseline84%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks terminal-bench/terminal-bench-2-1 --batch_size auto Benchmark Specs