Best AI Models for Reasoning

The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).

updated Oct 9, 2026 · 15 models ranked
How this list is ranked: Ranked by the reasoning pillar using the available published results and the current snapshot’s benchmark weights.

Current inputs: arc agi 1 (weight 1); arc agi 2 (weight 1); arc agi 3 (weight 1); arc agi v1 public eval (weight 1); arc agi v1 semi private (weight 1); arc agi v2 public eval (weight 1); arc agi v2 semi private (weight 1); arc agi v3 semi private (weight 1); gpqa diamond (weight 0.5); hle (weight 1); hle 1811 verified and revised items (weight 1); hle full set (weight 1); hle full set text mm (weight 1); hle full set tools (weight 1); hle scale (weight 1); hle text only (weight 1); hle text only subset (weight 1); hle text only subset tools (weight 1); hle text only tools (weight 1); hle tools (weight 1); livebench reasoning connections 2026_06_25 (weight 1); livebench reasoning consecutive events 2026_06_25 (weight 1); livebench reasoning logic with navigation 2026_06_25 (weight 1); livebench reasoning spatial 2026_06_25 (weight 1); livebench reasoning theory of mind 2026_06_25 (weight 1); livebench reasoning zebra puzzle 2026_06_25 (weight 1); mmlu pro (weight 0.5). Pillar scores can have different evidence coverage; review the individual results.

# Model Reasoning pillar SI Score Confidence Price in Context
1 Claude Opus 5.5 Anthropic 91.4 78.4 100% confidence 100 percent, Full $4.00 1M
2 Muse Spark 1.3 Meta provisional 90.5 70.1 93% confidence 93 percent, High $1.25 1M
3 Muse Spark 1.2 Meta provisional 90.4 62.3 53% confidence 53 percent, Medium $1.25 1M
4 GLM-5.1 Z.ai provisional 89.9 63.3 64% confidence 64 percent, Medium $1.40 200K
5 Claude Fable 5 Anthropic provisional 89.4 76.8 100% confidence 100 percent, Full $10.00 1M
6 LongCat-2.0 meituan provisional 88.9 56.6 100% confidence 100 percent, Full — 1M
7 Claude Sonnet 5.5 Anthropic 88.6 69.4 93% confidence 93 percent, High $2.00 1M
8 GPT-6.1 Sol OpenAI 87.5 73.0 100% confidence 100 percent, Full $2.00 1.1M
9 Qwen3.6 Max Preview Alibaba / Qwen provisional 87.4 66.4 64% confidence 64 percent, Medium $1.30 262K
10 Claude Sonnet 5 Anthropic provisional 87.3 67.7 93% confidence 93 percent, High $2.00 1M
11 Claude Fable 5.1 Anthropic 87.1 80.2 100% confidence 100 percent, Full $10.00 1M
12 GPT-6 Astra OpenAI 87.1 77.6 100% confidence 100 percent, Full $10.00 1.1M
13 MiMo-V2.5-Pro xiaomi provisional 86.6 62.6 88% confidence 88 percent, High $0.43 1M
14 Qwen3.5 397B-A17B Alibaba / Qwen provisional 86.1 64.0 69% confidence 69 percent, Medium $0.17 262K
15 GPT-5.5 Pro OpenAI 85.9 63.5 53% confidence 53 percent, Medium $30.00 1.1M

Task pages rank on one transparent metric, not the blended SI Score — see methodology for how pillars and confidence are computed. Missing values mean the source has not reported them.