Suite Selection Advisor

Benchmark Advisor

Select your target capability, evaluation priority, and model tier to find calibrated, non-saturated benchmark suites.

1Capability Domain
2Evaluation Priority
3Model Scale Under Test

Recommended Suites for Coding

Calibrated for Frontier Reasoning with focus on Frontier Headroom.

5 Recommended
#1FrontierCodeActiveFrontier Discriminator
Trust Score
95

Sufficiently difficult to discriminate frontier reasoning models with test-time compute.

Expected Score50% – 95%
ContaminationLOW
Human Baseline45%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks frontiercode --batch_size auto
Benchmark Specs
#2LiveCodeBenchSaturatedFrontier Discriminator
Trust Score
35

Sufficiently difficult to discriminate frontier reasoning models with test-time compute.

Expected Score50% – 95%
ContaminationMEDIUM
Human Baseline65%
Refresh Cadencecontinuous
lm_eval --model hf --model_args pretrained=<your_model> --tasks --scenario codegeneration --batch_size auto
Benchmark Specs
#3SWE-bench ProActiveFrontier Discriminator
Trust Score
90

Sufficiently difficult to discriminate frontier reasoning models with test-time compute.

Expected Score50% – 95%
ContaminationMEDIUM
Human Baseline74%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks swe_bench_pro_eval.py --batch_size auto
Benchmark Specs
#4SWE-bench VerifiedSaturatedFrontier Discriminator
Trust Score
25

Sufficiently difficult to discriminate frontier reasoning models with test-time compute.

Expected Score50% – 95%
ContaminationHIGH
Human Baseline78.4%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks SWE-bench_Verified --batch_size auto
Benchmark Specs
#5Terminal-BenchActiveRecommended Suite
Trust Score
80

Active benchmark with solid headroom against top foundation models.

Expected Score70% – 90%
ContaminationHIGH
Human Baseline84%
Refresh Cadencestatic
lm_eval --model hf --model_args pretrained=<your_model> --tasks terminal-bench/terminal-bench-2-1 --batch_size auto
Benchmark Specs