Best AI Models for Math

Mathematical-reasoning models ordered by the math pillar from the available published evaluations.

updated Oct 9, 2026 · 15 models ranked
How this list is ranked: Ranked by the math pillar using the snapshot’s benchmark weights. Check coverage and individual evaluations before treating a small lead as decisive.

Current inputs: frontiermath tier 1 3 (weight 1); frontiermath tier 4 (weight 1); frontiermath tier 4 v2 (weight 1); frontiermath tiers 1 3 v2 (weight 1); frontiermath vv2 tier 1 3 (weight 1); frontiermath vv2 tier 4 (weight 1); livebench math amps hard 2026_06_25 (weight 1); livebench math integrals with game 2026_06_25 (weight 1); livebench math math comp 2026_06_25 (weight 1); livebench math olympiad 2026_06_25 (weight 1); livebench math simplify 2026_06_25 (weight 1); math level 5 (weight 1); otis mock aime 2024 2025 (weight 0.5). Pillar scores can have different evidence coverage; review the individual results.

# Model Math pillar SI Score Confidence Price in Context
1 GPT-6.1 Sol OpenAI 93.0 73.0 100% confidence 100 percent, Full $2.00 1.1M
2 GPT-6 Astra OpenAI 92.9 77.6 100% confidence 100 percent, Full $10.00 1.1M
3 Claude Opus 5.5 Anthropic 91.5 78.4 100% confidence 100 percent, Full $4.00 1M
4 Claude Fable 5.1 Anthropic 91.5 80.2 100% confidence 100 percent, Full $10.00 1M
5 Gemini 3 Pro Preview Google provisional 91.4 58.6 68% confidence 68 percent, Medium — 1M
6 GPT-6 Sol OpenAI 91.3 68.9 100% confidence 100 percent, Full $2.00 1.1M
7 Qwen3.6 Max Preview Alibaba / Qwen provisional 91.1 66.4 64% confidence 64 percent, Medium $1.30 262K
8 Claude Fable 5 Anthropic provisional 91.0 76.8 100% confidence 100 percent, Full $10.00 1M
9 Claude Sonnet 5.5 Anthropic 90.5 69.4 93% confidence 93 percent, High $2.00 1M
10 GPT-5.6 Sol OpenAI provisional 89.3 72.1 100% confidence 100 percent, Full $4.00 1.1M
11 GPT OSS 120B OpenAI provisional 88.9 54.8 79% confidence 79 percent, Medium — 131K
12 DeepSeek V4.1 Flash DeepSeek 88.0 66.2 69% confidence 69 percent, Medium $0.15 1M
13 Claude Opus 5 Anthropic provisional 86.9 74.9 100% confidence 100 percent, Full $5.00 1M
14 Muse Spark 1.2 Meta provisional 86.7 62.3 53% confidence 53 percent, Medium $1.25 1M
15 DeepSeek-R1 DeepSeek provisional 86.6 47.9 88% confidence 88 percent, High $0.55 128K

Task pages rank on one transparent metric, not the blended SI Score — see methodology for how pillars and confidence are computed. Missing values mean the source has not reported them.