FireBench
Evaluates instruction following for enterprise/API pipelines — exact output format, strict step ordering, ranking, and calibrated refusal across six categories.
Newly released (arXiv paper submitted 5 March 2026); no refresh cycle or successor yet, and scores still separate frontier models. Worth tracking but with no independent replications published so far.
Performance Timeline
Longitudinal progression of model scores against human baselines.Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human baseline on fine-grained instruction compliance and negative constraint adherence.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.composite instruction-following scoreDataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks firebench --batch_size autoopencompass --datasets firebench --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for FireBench, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace