Benchmark Index
Every number has a source

AI model benchmarks, with receipts.

Frontier-model benchmark results aggregated from vendor announcements and leaderboards — source-linked, verification-labelled, and honest about what the numbers can't tell you. How we verify scores. Last reviewed 2026-07-21.

SWE-bench Verified leaderboard

Resolving real GitHub issues from real repositories; human-validated subset of SWE-bench.

Claude Fable 5: sourceGPT-5.5: sourceClaude Opus 4.8: sourceGemini 3.1 Pro: sourceDeepSeek V4: source

Claude Sonnet 5: not yet verified — never plotted (Placeholder — demonstrates the 'not yet verified' rendering state.)

Most scores are vendor-reported and scaffold-dependent; results vary by ±2–3 points across harnesses.

Compare models across benchmarks

Pick up to three models.

Missing bars mean no sourced score for that model × benchmark. Cross-vendor caveats apply — see the methodology.

Coding ability vs. price

SWE-bench Verified score against output price (USD per million tokens, log scale). Only models with sourced pricing appear.

Claude Fable 5 ($50/M src)Claude Opus 4.8 ($25/M src)GPT-5.5 ($30/M src)Gemini 3.1 Pro ($12/M src)DeepSeek V4 ($0.87/M src)

Pricing not yet verified for: GPT-5.6 — excluded rather than guessed.

Recent releases

Notable frontier-model releases, source-linked. Updated editorially.

  1. 2026-06-09

    Claude Fable 5 (Mythos-class) · Anthropic

    First generally available model of Anthropic's new Mythos tier above Opus; vendor-reported 95% on SWE-bench Verified and 80.3% on SWE-bench Pro. announcement

  2. 2026-04-23

    GPT-5.5 · OpenAI

    Large jump on abstract reasoning (85% ARC-AGI-2, +11.7 points over GPT-5.4) and strong agentic computer-use results. announcement

  3. ~2026-03 (approximate)

    Gemini 3.1 Pro · Google DeepMind

    Led 13 of 16 vendor-evaluated benchmarks at launch, including GPQA Diamond (94.3%) and the no-tools Humanity's Last Exam (44.4%). announcement

  4. ~2026-01 (approximate)

    GPT-5.2 · OpenAI

    Reported as the first model above 90% on ARC-AGI-1, shifting expectations for inference-time reasoning budgets. announcement

All scores

Every entry with its verification level and source. We never publish a number without one.

ModelBenchmarkScoreVerificationSource
Claude Fable 5SWE-bench Verified95.0%
secondary source
source
Claude Fable 5SWE-bench Pro80.3%
secondary source
source
Claude Fable 5GPQA Diamond92.6%
secondary source
source
Claude Opus 4.8SWE-bench Verified88.6%
secondary source
source
Claude Opus 4.8SWE-bench Pro69.2%
secondary source
source
Claude Opus 4.8GPQA Diamond93.6%
secondary source
source
GPT-5.5SWE-bench Verified88.7%
secondary source
source
GPT-5.5SWE-bench Pro58.6%
secondary source
source
GPT-5.5GPQA Diamond93.6%
secondary source
source
GPT-5.5ARC-AGI-285.0%
secondary source
source
GPT-5.5OSWorld-Verified78.7%
secondary source
source
GPT-5.5FrontierMath (Tier 4)39.6%
secondary source
source
Gemini 3.1 ProSWE-bench Verified80.6%
secondary source
source
Gemini 3.1 ProSWE-bench Pro54.2%
secondary source
source
Gemini 3.1 ProGPQA Diamond94.3%
secondary source
source
Gemini 3.1 ProARC-AGI-277.1%
secondary source
source
Gemini 3.1 ProHumanity's Last Exam44.4%
secondary source
source
DeepSeek V4
open
SWE-bench Verified80.6%
secondary source
source
Kimi K2.6
open
SWE-bench Pro58.6%
secondary source
source
Claude Sonnet 5SWE-bench Verifiednot yet verified
not yet verified
GPT-5.6ARC-AGI-2not yet verified
not yet verified
source