Best AI Models for Reasoning
The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).
Current inputs: arc agi 1 (weight 1); arc agi 2 (weight 1); arc agi 3 (weight 1); arc agi v1 public eval (weight 1); arc agi v1 semi private (weight 1); arc agi v2 public eval (weight 1); arc agi v2 semi private (weight 1); arc agi v3 semi private (weight 1); gpqa diamond (weight 0.5); hle (weight 1); hle 1811 verified and revised items (weight 1); hle full set (weight 1); hle full set text mm (weight 1); hle full set tools (weight 1); hle scale (weight 1); hle text only (weight 1); hle text only subset (weight 1); hle text only subset tools (weight 1); hle text only tools (weight 1); hle tools (weight 1); livebench reasoning connections 2026_06_25 (weight 1); livebench reasoning consecutive events 2026_06_25 (weight 1); livebench reasoning logic with navigation 2026_06_25 (weight 1); livebench reasoning spatial 2026_06_25 (weight 1); livebench reasoning theory of mind 2026_06_25 (weight 1); livebench reasoning zebra puzzle 2026_06_25 (weight 1); mmlu pro (weight 0.5). Pillar scores can have different evidence coverage; review the individual results.
| # | Model | Reasoning pillar | SI Score | Confidence |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 91.4 | 78.4 | 100% confidence 100 percent, Full |
| 2 | Muse Spark 1.3 Meta provisional | 90.5 | 70.1 | 93% confidence 93 percent, High |
| 3 | Muse Spark 1.2 Meta provisional | 90.4 | 62.3 | 53% confidence 53 percent, Medium |
| 4 | GLM-5.1 Z.ai provisional | 89.9 | 63.3 | 64% confidence 64 percent, Medium |
| 5 | Claude Fable 5 Anthropic provisional | 89.4 | 76.8 | 100% confidence 100 percent, Full |
| 6 | LongCat-2.0 meituan provisional | 88.9 | 56.6 | 100% confidence 100 percent, Full |
| 7 | Claude Sonnet 5.5 Anthropic | 88.6 | 69.4 | 93% confidence 93 percent, High |
| 8 | GPT-6.1 Sol OpenAI | 87.5 | 73.0 | 100% confidence 100 percent, Full |
| 9 | Qwen3.6 Max Preview Alibaba / Qwen provisional | 87.4 | 66.4 | 64% confidence 64 percent, Medium |
| 10 | Claude Sonnet 5 Anthropic provisional | 87.3 | 67.7 | 93% confidence 93 percent, High |
| 11 | Claude Fable 5.1 Anthropic | 87.1 | 80.2 | 100% confidence 100 percent, Full |
| 12 | GPT-6 Astra OpenAI | 87.1 | 77.6 | 100% confidence 100 percent, Full |
| 13 | MiMo-V2.5-Pro xiaomi provisional | 86.6 | 62.6 | 88% confidence 88 percent, High |
| 14 | Qwen3.5 397B-A17B Alibaba / Qwen provisional | 86.1 | 64.0 | 69% confidence 69 percent, Medium |
| 15 | GPT-5.5 Pro OpenAI | 85.9 | 63.5 | 53% confidence 53 percent, Medium |
Task pages rank on one transparent metric, not the blended SI Score — see methodology for how pillars and confidence are computed. Missing values mean the source has not reported them.