terminal bench v4 0 Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 Anthropic | 70.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 70.6 | 93% confidence 93 percent, High |
| 2 | Claude Opus 5.5 Anthropic | 66.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; production safeguards with fallback; 4.0Published Sep 22, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 66.4 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 64.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 64.8 | 100% confidence 100 percent, Full |
| 4 | Claude Sonnet 5.5 Anthropic | 61.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 61.8 | 93% confidence 93 percent, High |
| 5 | GPT-6 Astra OpenAI | 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 58.2 | 100% confidence 100 percent, Full |
| 6 | GPT-6.1 Sol OpenAI | 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 58.2 | 100% confidence 100 percent, Full |
| 7 | GPT-6 Astra OpenAI | 57.9%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 8 | Claude Fable 5.1 Anthropic | 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 9 | Claude Fable 5.1 Anthropic | 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 10 | GPT-6 Astra OpenAI | 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; highPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 11 | GPT-6 Astra OpenAI | 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhighPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 12 | Claude Fable 5.1 Anthropic | 55.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] production safeguards with fallback; 4.0Published Sep 1, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 55.8 | 100% confidence 100 percent, Full |
| 13 | Claude Fable 5.1 Anthropic | 54.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 54.5 | 100% confidence 100 percent, Full |
| 14 | GPT-6 Astra OpenAI | 54.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; mediumPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 54.2 | 100% confidence 100 percent, Full |
| 15 | Claude Fable 5.1 Anthropic | 53.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 53.9 | 100% confidence 100 percent, Full |
| 16 | Claude Opus 5 Anthropic | 53.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 53.9 | 100% confidence 100 percent, Full |
| 17 | Claude Opus 5 Anthropic | 51.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 51.8 | 100% confidence 100 percent, Full |
| 18 | GPT-6 Astra OpenAI | 50.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; lowPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 50.6 | 100% confidence 100 percent, Full |
| 19 | Claude Opus 5 Anthropic | 50.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 50.3 | 100% confidence 100 percent, Full |
| 20 | GPT-6 Sol OpenAI | 49.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 49.4 | 100% confidence 100 percent, Full |
| 21 | Claude Opus 5 Anthropic | 44.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 44.9 | 100% confidence 100 percent, Full |
| 22 | Claude Fable 5 Anthropic | 44.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 44.5 | 100% confidence 100 percent, Full |
| 23 | Claude Fable 5.1 Anthropic | 43.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 43.3 | 100% confidence 100 percent, Full |
| 24 | GLM-5.3 Z.ai | 41.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 41.8 | 93% confidence 93 percent, High |
| 25 | Claude Haiku 5.5 Anthropic | 39.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 39.2 | 100% confidence 100 percent, Full |
| 26 | Grok 4.7 xAI | 37.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; 4.0Published Sep 21, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 37.6 | 100% confidence 100 percent, Full |
| 27 | Grok 4.7 xAI | 37.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; xhighPublished Sep 21, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 37.6 | 100% confidence 100 percent, Full |
| 28 | GPT-5.6 Sol OpenAI | 37.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 37.3 | 100% confidence 100 percent, Full |
| 29 | GLM-5.3-Flash Z.ai | 35.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 35.8 | 100% confidence 100 percent, Full |
| 30 | MiMo-V2.6-Pro xiaomi | 34.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.6 Pro comparison column; 4.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 34.9 | 88% confidence 88 percent, High |
| 31 | Claude Opus 5 Anthropic | 34.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 34.9 | 100% confidence 100 percent, Full |
| 32 | DeepSeek V4.1 Flash DeepSeek | 31.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness minimal mode; 4.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 31.2 | 69% confidence 69 percent, Medium |
| 33 | MiMo-V2.6-Flash xiaomi | 28.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 28.8 | 88% confidence 88 percent, High |
| 34 | Qwen3.8 Max 0902 Alibaba / Qwen | 27.0%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 27.0 | 40% confidence 40 percent, Low |
| 35 | Claude Opus 4.8 Anthropic | 23.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 23.6 | 100% confidence 100 percent, Full |
| 36 | GPT-5.6 Terra OpenAI | 21.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 21.5 | 100% confidence 100 percent, Full |
| 37 | Grok 4.6 xAI | 20.3%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] high effort; 4.0Published Sep 21, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 20.3 | 100% confidence 100 percent, Full |
| 38 | Grok 4.6 xAI | 20.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; highPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 20.3 | 100% confidence 100 percent, Full |
| 39 | Gemini 3.8 Flash Google | 19.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 19.1 | 100% confidence 100 percent, Full |
| 40 | Gemini 3.8 Flash Google | 19.1%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] mini-SWE-agent; highPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 19.1 | 100% confidence 100 percent, Full |
| 41 | GPT-5.6 Luna OpenAI | 17.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 17.3 | 100% confidence 100 percent, Full |
| 42 | GPT-6 Luna OpenAI | 16.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 16.4 | 100% confidence 100 percent, Full |
| 43 | Muse Spark 1.3 Meta | 14.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Muse Code; xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 14.6 | 93% confidence 93 percent, High |
| 44 | Claude Sonnet 5 Anthropic | 12.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 12.4 | 93% confidence 93 percent, High |
| 45 | Grok 4.5 xAI | 12.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; highPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 12.4 | 100% confidence 100 percent, Full |
| 46 | Gemini 3.7 Flash Google | 11.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] mini-SWE-agent; highPublished Sep 3, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ | 11.2 | 100% confidence 100 percent, Full |
| 47 | MiMo-V2.5-Pro xiaomi | 1.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.5 Pro comparison column; 4.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 1.5 | 88% confidence 88 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.