GPT-6 Astra vs Claude Fable 5
VS
SI Score, rank and confidence side by side
Pillars — GPT-6 Astra
Coding (weight 40 percent) 64.8 / 70.2
Math (weight 15 percent) 92.9 / 91.0
Preference (weight 15 percent) 77.1 / 81.1
Reasoning (weight 30 percent) 87.1 / 89.4
Right-hand value is Claude Fable 5.
The basics
| Attribute | GPT-6 Astra | Claude Fable 5 |
|---|---|---|
| Input price / 1M | $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $10.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Output price / 1M | $50.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $50.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Context window | 1.1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ | 1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Released | Sep 3, 2026OpenAI API changelogPublished source fact
Retrieved Oct 9, 2026 · factual citation Open source ↗ | Jun 9, 2026models.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Open weights | Closed | Closed |
Shared benchmarks
27 in common| Benchmark | GPT-6 Astra | Claude Fable 5 | Conditions |
|---|---|---|---|
| arc agi v1 public eval | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-high 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-low 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-max 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-medium 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhigh | Check variant and harness |
| arc agi v1 semi private | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-high 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-low 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-max 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-medium 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhigh | Check variant and harness |
| arc agi v2 public eval | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 96.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 93.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-high 80.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-low 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-max 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-medium 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhigh | Check variant and harness |
| arc agi v2 semi private | 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-high 76.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-low 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-max 82.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-medium 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhigh | Check variant and harness |
| frontiermath tier 4 v2 | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] high 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] medium 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] none 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhigh | 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| frontiermath tiers 1 3 v2 | 93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 87.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| gpqa diamond | 95.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max 96%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] high 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| hle tools | 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools | 64.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools | Check variant and harness |
| livebench coding code completion 2026_06_25 | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding code generation 2026_06_25 | 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 91.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding javascript 2026_06_25 | 63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding python 2026_06_25 | 65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding typescript 2026_06_25 | 43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 53.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math amps hard 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math integrals with game 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math math comp 2026_06_25 | 97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math olympiad 2026_06_25 | 92.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math simplify 2026_06_25 | 70.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 72.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning connections 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning consecutive events 2026_06_25 | 90.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 91.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning logic with navigation 2026_06_25 | 88%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 78%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning spatial 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 96%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning theory of mind 2026_06_25 | 84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning zebra puzzle 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| lmarena text | 1442.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | 1491.7 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | Check variant and harness |
| otis mock aime 2024 2025 | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] high 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 99.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| terminal bench v4 0 | 57.9%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; highPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; high 50.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; lowPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; low 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; max 54.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; mediumPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; medium 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhighPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhigh | 44.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; max | Check variant and harness |
Only benchmarks both models have are compared. Raw values are in each benchmark's own unit; a direct comparison requires matching evaluation conditions. All reported source/variant rows are shown; no raw-result winner is assigned across unmatched harnesses. Sources and dates sit behind every dotted number.
Which should you choose?
What the data says — heuristics from the numbers above, not a verdict:
- For coding, Claude Fable 5 leads by 5.4 normalized points.
- For preference, Claude Fable 5 leads by 4.1 normalized points.
- Confidence differs: GPT-6 Astra 100% vs Claude Fable 5 100% — pending sources can still move either score.