AI model benchmarks, with receipts.
Frontier-model benchmark results aggregated from vendor announcements and leaderboards — source-linked, verification-labelled, and honest about what the numbers can't tell you. How we verify scores. Last reviewed 2026-07-21.
Resolving real GitHub issues from real repositories; human-validated subset of SWE-bench.
Claude Sonnet 5: not yet verified — never plotted (Placeholder — demonstrates the 'not yet verified' rendering state.)
Most scores are vendor-reported and scaffold-dependent; results vary by ±2–3 points across harnesses.
Pick up to three models.
Missing bars mean no sourced score for that model × benchmark. Cross-vendor caveats apply — see the methodology.
SWE-bench Verified score against output price (USD per million tokens, log scale). Only models with sourced pricing appear.
Notable frontier-model releases, source-linked. Updated editorially.
2026-06-09
Claude Fable 5 (Mythos-class) · Anthropic
First generally available model of Anthropic's new Mythos tier above Opus; vendor-reported 95% on SWE-bench Verified and 80.3% on SWE-bench Pro. announcement
2026-04-23
GPT-5.5 · OpenAI
Large jump on abstract reasoning (85% ARC-AGI-2, +11.7 points over GPT-5.4) and strong agentic computer-use results. announcement
~2026-03 (approximate)
Gemini 3.1 Pro · Google DeepMind
Led 13 of 16 vendor-evaluated benchmarks at launch, including GPQA Diamond (94.3%) and the no-tools Humanity's Last Exam (44.4%). announcement
~2026-01 (approximate)
GPT-5.2 · OpenAI
Reported as the first model above 90% on ARC-AGI-1, shifting expectations for inference-time reasoning budgets. announcement
Every entry with its verification level and source. We never publish a number without one.
| Model | Benchmark | Score | Verification | Source |
|---|---|---|---|---|
| Claude Fable 5 | SWE-bench Verified | 95.0% | secondary source | source |
| Claude Fable 5 | SWE-bench Pro | 80.3% | secondary source | source |
| Claude Fable 5 | GPQA Diamond | 92.6% | secondary source | source |
| Claude Opus 4.8 | SWE-bench Verified | 88.6% | secondary source | source |
| Claude Opus 4.8 | SWE-bench Pro | 69.2% | secondary source | source |
| Claude Opus 4.8 | GPQA Diamond | 93.6% | secondary source | source |
| GPT-5.5 | SWE-bench Verified | 88.7% | secondary source | source |
| GPT-5.5 | SWE-bench Pro | 58.6% | secondary source | source |
| GPT-5.5 | GPQA Diamond | 93.6% | secondary source | source |
| GPT-5.5 | ARC-AGI-2 | 85.0% | secondary source | source |
| GPT-5.5 | OSWorld-Verified | 78.7% | secondary source | source |
| GPT-5.5 | FrontierMath (Tier 4) | 39.6% | secondary source | source |
| Gemini 3.1 Pro | SWE-bench Verified | 80.6% | secondary source | source |
| Gemini 3.1 Pro | SWE-bench Pro | 54.2% | secondary source | source |
| Gemini 3.1 Pro | GPQA Diamond | 94.3% | secondary source | source |
| Gemini 3.1 Pro | ARC-AGI-2 | 77.1% | secondary source | source |
| Gemini 3.1 Pro | Humanity's Last Exam | 44.4% | secondary source | source |
| DeepSeek V4 open | SWE-bench Verified | 80.6% | secondary source | source |
| Kimi K2.6 open | SWE-bench Pro | 58.6% | secondary source | source |
| Claude Sonnet 5 | SWE-bench Verified | not yet verified | not yet verified | — |
| GPT-5.6 | ARC-AGI-2 | not yet verified | not yet verified | source |