RULER
Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.
RULER is still a current synthetic stress test, but newer top rows report mid-90s average scores across 4k-128k, so it is approaching ceiling for the strongest long-context models while still exposing degradation.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| Jamba-1.5-large (README Avg. snapshot) | 96% | 2024-08-01 | independent | Source ↗ |
| Gemini-1.5-pro (README Avg. snapshot) | 95.8% | 2024-08-01 | independent | Source ↗ |
| Qwen2.5-14B-Instruct-1M (README Avg. snapshot) | 95.7% | 2024-08-01 | independent | Source ↗ |
| Qwen3-235B-A22B (README Avg. snapshot) | 95% | 2024-08-01 | independent | Source ↗ |
| Qwen3-14B (README Avg. snapshot) | 94.6% | 2024-08-01 | independent | Source ↗ |
| Jamba-1.5-mini (README Avg. snapshot) | 93.9% | 2024-08-01 | independent | Source ↗ |
| Qwen3-32B (README Avg. snapshot) | 93.7% | 2024-08-01 | independent | Source ↗ |
| EXAONE-4.0-32B (README Avg. snapshot) | 93.1% | 2024-08-01 | independent | Source ↗ |
| Qwen2.5-7B-Instruct-1M (README Avg. snapshot) | 91.8% | 2024-08-01 | independent | Source ↗ |
| GPT-4-1106-preview (README Avg. snapshot) | 91.6% | 2024-08-01 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human accuracy on complex synthetic key-value, aggregation, and tracing stress tests at 128K context.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.accuracy across configured synthetic tasksDataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks ruler --batch_size autoopencompass --datasets ruler --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for RULER, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace