Methodology
Benchmark Index aggregates published AI benchmark results. We do not run evaluations ourselves (yet); we collect, source and contextualize numbers others report.
Sourcing rules
- Every score links to its source: a vendor model card, an official leaderboard, or press coverage of either.
- Primary source means we checked the number against an official vendor or benchmark-organisation page.
- Secondary source means the number comes from reputable coverage or an aggregator and has not yet been checked against the original.
- Not yet verified entries show no number at all — we never estimate or interpolate scores.
Pricing data
Model prices (USD per million tokens) follow the same rules: Anthropic prices come from the official pricing documentation (primary); other vendors' prices come from reputable aggregators (secondary) until checked against the vendor's own pricing page. Models without a sourced price are excluded from price charts rather than estimated.
Why scores disagree
- Vendor scaffolding: agentic benchmarks (SWE-bench and similar) depend heavily on the harness around the model.
- Compute budget: reasoning benchmarks vary with thinking-time and pass@k settings.
- Tools versus no-tools: some benchmarks (for example Humanity's Last Exam) report both variants and they differ widely.
- Run-to-run noise: near-saturated benchmarks make one-point gaps meaningless.
Each benchmark page carries its own caveats. Comparisons are presented vendor-neutrally; we have no affiliate relationships with any model vendor.
Corrections
Spotted an error or a fresher source? Reply to any newsletter issue — every entry shows its last-reviewed date and we correct fast.