Best AI Models for Coding
Coding and agentic models ordered by their published coding pillar, combining the benchmark variants available in this snapshot.
Current inputs: aider polyglot (weight 1); livebench coding code completion 2026_06_25 (weight 1); livebench coding code generation 2026_06_25 (weight 1); livebench coding javascript 2026_06_25 (weight 1); livebench coding python 2026_06_25 (weight 1); livebench coding typescript 2026_06_25 (weight 1); swe bench pro (weight 1); swe bench pro public (weight 1); swe bench pro qwen refined and corrected task set (weight 1); swe bench pro system card 2026 06 (weight 1); swe bench verified (weight 1); terminal bench (weight 1); terminal bench v0 1 (weight 1); terminal bench v2 0 (weight 1); terminal bench v2 1 (weight 1); terminal bench v3 0 (weight 1); terminal bench v4 0 (weight 1). Pillar scores can have different evidence coverage; review the individual results.
| # | Model | Coding pillar | SI Score | Confidence |
|---|---|---|---|---|
| 1 | Fugu Ultra sakana provisional | 77.9 | 57.9 | 100% confidence 100 percent, Full |
| 2 | Qwen3.6 Max Preview Alibaba / Qwen provisional | 76.7 | 66.4 | 64% confidence 64 percent, Medium |
| 3 | Qwen3.5 397B-A17B Alibaba / Qwen provisional | 76.4 | 64.0 | 69% confidence 69 percent, Medium |
| 4 | MiniMax-M2.5 minimax provisional | 75.8 | 57.8 | 100% confidence 100 percent, Full |
| 5 | GLM-5.1 Z.ai provisional | 74.2 | 63.3 | 64% confidence 64 percent, Medium |
| 6 | Claude Opus 4.1 Anthropic provisional | 73.3 | 51.5 | 80% confidence 80 percent, High |
| 7 | DeepSeek V4.1 Flash DeepSeek | 73.3 | 66.2 | 69% confidence 69 percent, Medium |
| 8 | Claude Opus 5.5 Anthropic | 72.5 | 78.4 | 100% confidence 100 percent, Full |
| 9 | GLM-5 Z.ai provisional | 72.4 | 59.4 | 90% confidence 90 percent, High |
| 10 | Kimi K2.5 Moonshot AI provisional | 72.4 | 59.0 | 100% confidence 100 percent, Full |
| 11 | Claude Sonnet 4.5 Anthropic provisional | 71.3 | 53.9 | 80% confidence 80 percent, High |
| 12 | Kimi K3 Moonshot AI provisional | 71.1 | 71.0 | 100% confidence 100 percent, Full |
| 13 | Claude Opus 4 Anthropic provisional | 70.6 | 55.5 | 100% confidence 100 percent, Full |
| 14 | Claude Fable 5 Anthropic provisional | 70.2 | 76.8 | 100% confidence 100 percent, Full |
| 15 | DeepSeek V3.2 DeepSeek provisional | 70.0 | 53.3 | 56% confidence 56 percent, Medium |
Task pages rank on one transparent metric, not the blended SI Score — see methodology for how pillars and confidence are computed. Missing values mean the source has not reported them.