Claude Opus 5.5 vs Claude Opus 5
VS
SI Score, rank and confidence side by side
Pillars — Claude Opus 5.5
Coding (weight 40 percent) 72.5 / 69.7
Math (weight 15 percent) 91.5 / 86.9
Preference (weight 15 percent) 82.6 / 81.9
Reasoning (weight 30 percent) 91.4 / 84.3
Right-hand value is Claude Opus 5.
The basics
| Attribute | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Input price / 1M | $4.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $5.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Output price / 1M | $20.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $25.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Context window | 1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Released | Sep 22, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | Jul 24, 2026models.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Open weights | Closed | Closed |
Shared benchmarks
27 in common| Benchmark | Claude Opus 5.5 | Claude Opus 5 | Conditions |
|---|---|---|---|
| arc agi v1 public eval | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-high 92.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-low 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-max 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-medium 98.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhigh | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-high 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-max | Check variant and harness |
| arc agi v1 semi private | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-high 88.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-low 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-max 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-medium 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhigh | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-high 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-max | Check variant and harness |
| arc agi v2 public eval | 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-high 67.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-low 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-max 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-medium 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhigh | 93.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-high 97.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-max | Check variant and harness |
| arc agi v2 semi private | 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-high 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-low 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-max 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-medium 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhigh | 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-high 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-max | Check variant and harness |
| frontiermath tier 4 v2 | 95%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| frontiermath tiers 1 3 v2 | 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| gpqa diamond | 90.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 92.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| hle tools | 67.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools; production safeguards with fallbackPublished Sep 22, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools; production safeguards with fallback | 64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools | Check variant and harness |
| livebench coding code completion 2026_06_25 | 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding code generation 2026_06_25 | 91.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding javascript 2026_06_25 | 72.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding python 2026_06_25 | 70%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 75%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding typescript 2026_06_25 | 53.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math amps hard 2026_06_25 | 99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 99.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math integrals with game 2026_06_25 | 99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math math comp 2026_06_25 | 97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 94.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math olympiad 2026_06_25 | 92.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math simplify 2026_06_25 | 63%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 61.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning connections 2026_06_25 | 99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning consecutive events 2026_06_25 | 90.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 77.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning logic with navigation 2026_06_25 | 80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 86%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning spatial 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning theory of mind 2026_06_25 | 84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning zebra puzzle 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| lmarena text | 1511.7 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | 1502.1 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | Check variant and harness |
| otis mock aime 2024 2025 | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| terminal bench v4 0 | 66.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; production safeguards with fallback; 4.0Published Sep 22, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; production safeguards with fallback; 4.0 64.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; max | 50.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; high 34.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; low 51.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; max 44.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; medium 53.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhigh | Check variant and harness |
Only benchmarks both models have are compared. Raw values are in each benchmark's own unit; a direct comparison requires matching evaluation conditions. All reported source/variant rows are shown; no raw-result winner is assigned across unmatched harnesses. Sources and dates sit behind every dotted number.
Which should you choose?
What the data says — heuristics from the numbers above, not a verdict:
- For reasoning, Claude Opus 5.5 leads by 7.0 normalized points.
- For math, Claude Opus 5.5 leads by 4.6 normalized points.
- Claude Opus 5.5 has the lower input-token price ($4.00 vs $5.00 per 1M).
- Confidence differs: Claude Opus 5.5 100% vs Claude Opus 5 100% — pending sources can still move either score.