SI Score · updated Oct 8, 2026

Superintelligence leaderboard.
Open evidence. Clear rankings.

The latest frontier models, ranked by the SI Score — a composite of open benchmarks, scaled per benchmark and weighted across reasoning, coding, math and human preference. Every number links to its source and date, and each score carries a confidence % that rises as sources report.

447
models tracked
57
benchmarks in the composite
22
sources tracked
80.2
top SI Score — Claude Fable 5.1

The leaderboard

Full catalog →
1 Claude Fable 5.1 Anthropic 80.2 100% confidence 100 percent, Full $10.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
$50.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
Sep 1, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
2 Claude Opus 5.5 Anthropic 78.4 100% confidence 100 percent, Full $4.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
$20.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
Sep 22, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
3 GPT-6 Astra OpenAI 77.6 100% confidence 100 percent, Full $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
$50.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 3, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
4 Claude Fable 5 Anthropic 76.8 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 9, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
5 Claude Opus 5 Anthropic 74.9 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 24, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
6 GPT-6.1 Sol OpenAI 73.0 100% confidence 100 percent, Full $2.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
$10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 29, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
7 GPT-5.6 Sol OpenAI 72.1 100% confidence 100 percent, Full not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 9, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
8 Kimi K3 Moonshot AI 71.0 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 16, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
9 Claude Opus 4.6 Anthropic 70.5 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 5, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
10 Muse Spark 1.3 Meta 70.1 93% confidence 93 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 2, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
11 Claude Opus 4.7 Anthropic 70.1 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 16, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
12 Gemini 3.7 Flash Google 70.0 100% confidence 100 percent, Full $0.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$3.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 13, 2026Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
13 GPT-5.5 OpenAI 69.5 100% confidence 100 percent, Full not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 24, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
14 Claude Sonnet 5.5 Anthropic 69.4 93% confidence 93 percent, High $2.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
$10.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
Sep 28, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
15 GPT-6 Sol OpenAI 68.9 100% confidence 100 percent, Full not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 22, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
16 Gemini 3.8 Flash Google 68.8 100% confidence 100 percent, Full $0.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$3.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 2, 2026Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
17 Gemini 3.5 Flash Google 68.3 100% confidence 100 percent, Full $1.50Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$9.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 19, 2026Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
18 Claude Sonnet 4.6 Anthropic 68.2 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 17, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
19 GPT-5.6 Terra OpenAI 68.2 100% confidence 100 percent, Full not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 9, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
20 Gemini 3.1 Pro Preview Google 68.1 100% confidence 100 percent, Full $2.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$12.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 19, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
21 MiniMax-M3 minimax 68.0 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 1, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
22 GLM-5.3 Z.ai 67.7 93% confidence 93 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 14, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
23 Claude Sonnet 5 Anthropic 67.7 93% confidence 93 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 30, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
24 Claude Opus 4.8 Anthropic 67.5 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 28, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
25 Grok 4.6 xAI 66.6 100% confidence 100 percent, Full not yet reported not yet reported 500Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 12, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
26 Qwen3.6 Max Preview Alibaba / Qwen 66.4 64% confidence 64 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 20, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
27 DeepSeek V4.1 Flash DeepSeek 66.2 69% confidence 69 percent, Medium $0.15DeepSeek pricingOfficial off-peak uncached rate; peak is 2x; time schedule at source Retrieved Oct 8, 2026 · factual citation
Open source ↗
$0.60DeepSeek pricingOfficial off-peak uncached rate; peak is 2x; time schedule at source Retrieved Oct 8, 2026 · factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 10, 2026DeepSeek V4.1 Flash announcementOfficial dated introduction and availability announcement Retrieved Oct 8, 2026 · factual citation
Open source ↗
Open
28 GLM-5.3-Flash Z.ai 66.2 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 26, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
29 Kimi K2 Thinking Turbo Moonshot AI 66.0 64% confidence 64 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Nov 6, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
30 GLM-5.2 Z.ai 65.7 100% confidence 100 percent, Full not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 13, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
31 Gemini 3.6 Flash Google 65.6 100% confidence 100 percent, Full $0.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$3.75Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 21, 2026Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
32 Grok 4.7 xAI 65.5 100% confidence 100 percent, Full $2.00xAI models & pricingPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
$6.00xAI models & pricingPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
500KxAI models & pricingPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
Sep 21, 2026xAI models & pricingOfficial featured model page datePublished; model introduction date Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
33 Grok 4.5 xAI 64.8 100% confidence 100 percent, Full not yet reported not yet reported 500Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 8, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
34 Claude Haiku 5.5 Anthropic 64.8 100% confidence 100 percent, Full $0.10Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
$0.50Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
Oct 7, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
35 Claude Opus 4.5 Anthropic 64.6 100% confidence 100 percent, Full not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Nov 1, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
36 Kimi K2.6 Moonshot AI 64.0 85% confidence 85 percent, High not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 21, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
37 Qwen3.5 397B-A17B Alibaba / Qwen 64.0 69% confidence 69 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 15, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
38 GPT-5.5 Pro OpenAI 63.5 53% confidence 53 percent, Medium not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 24, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
39 GPT-5.6 Luna OpenAI 63.4 100% confidence 100 percent, Full not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 9, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
40 Gemma 4 26B A4B IT Google 63.4 64% confidence 64 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 2, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
41 GLM-5.1 Z.ai 63.3 64% confidence 64 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 7, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
42 DeepSeek V4 Pro DeepSeek 63.1 85% confidence 85 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 24, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
43 Gemma 4 31B IT Google 63.0 64% confidence 64 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 2, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
44 DeepSeek V4 Pro 0813 DeepSeek 63.0 69% confidence 69 percent, Medium $0.66DeepSeek pricingOfficial off-peak uncached rate; peak is 2x; time schedule at source Retrieved Oct 8, 2026 · factual citation
Open source ↗
$1.98DeepSeek pricingOfficial off-peak uncached rate; peak is 2x; time schedule at source Retrieved Oct 8, 2026 · factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 12, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
45 Qwen3.7 Plus Alibaba / Qwen 62.7 64% confidence 64 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 2, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
46 Qwen3.5 35B-A3B Alibaba / Qwen 62.7 64% confidence 64 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 23, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
47 MiMo-V2.5-Pro xiaomi 62.6 88% confidence 88 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 22, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
48 Muse Spark 1.2 Meta 62.3 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 5, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
49 Qwen3.8 Max Alibaba / Qwen 62.0 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 3, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
50 GPT-5.4 OpenAI 61.9 69% confidence 69 percent, Medium not yet reported not yet reported 1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 5, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
51 Qwen3.8 27B Alibaba / Qwen 61.9 69% confidence 69 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 14, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
52 GPT-6 Luna OpenAI 61.6 100% confidence 100 percent, Full $0.10OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
$0.50OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 8, 2026 · factual citation
Open source ↗
1.1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 22, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
53 Qwen3.6 Plus Alibaba / Qwen 61.6 80% confidence 80 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 2, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
54 GLM-4.7 Z.ai 61.3 69% confidence 69 percent, Medium not yet reported not yet reported 205Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 22, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
55 Muse Spark 1.1 Meta 61.2 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 8, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
56 DeepSeek V4 Flash 0731 DeepSeek 61.0 69% confidence 69 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 31, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
57 Hy3 tencent 60.7 88% confidence 88 percent, High not yet reported not yet reported 256Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 6, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
58 Qwen3.5 Flash Alibaba / Qwen 59.5 64% confidence 64 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 23, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
59 GLM-5 Z.ai 58.9 85% confidence 85 percent, High not yet reported not yet reported 205Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 12, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
60 Kimi K2.5 Moonshot AI 58.8 100% confidence 100 percent, Full not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jan 1, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
61 Qwen3.7 Max Alibaba / Qwen 58.6 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 21, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
62 GPT-5.5 Instant OpenAI 58.3 64% confidence 64 percent, Medium not yet reported not yet reported 400Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 5, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
63 Claude Haiku 4.5 Anthropic 58.1 64% confidence 64 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Oct 15, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
64 MiniMax-M2.7 minimax 57.9 88% confidence 88 percent, High not yet reported not yet reported 205Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 18, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
65 Grok 4.20 (Reasoning) xAI 57.7 80% confidence 80 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 9, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
66 MiMo-V2.6-Pro xiaomi 57.4 88% confidence 88 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 22, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
67 Gemini 3 Pro Preview Google 56.8 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Nov 18, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
68 MiniMax-M2.5 minimax 56.7 88% confidence 88 percent, High not yet reported not yet reported 205Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 12, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
69 GPT-5.4 nano OpenAI 56.6 69% confidence 69 percent, Medium not yet reported not yet reported 400Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 17, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
70 MiMo-V2.5 xiaomi 56.4 88% confidence 88 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 22, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
71 DeepSeek V4 Flash DeepSeek 56.2 53% confidence 53 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 24, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
72 GPT-5.4 mini OpenAI 55.8 69% confidence 69 percent, Medium not yet reported not yet reported 400Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 17, 2026OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
73 MiMo-V2.6-Flash xiaomi 55.8 88% confidence 88 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 22, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
74 MiniMax-M2 minimax 55.7 88% confidence 88 percent, High not yet reported not yet reported 205Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Oct 27, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
75 Qwen3.6 27B Alibaba / Qwen 55.6 53% confidence 53 percent, Medium not yet reported not yet reported 262Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 22, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
76 Gemini 3.5 Flash Lite Google 55.2 100% confidence 100 percent, Full $0.30Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$2.50Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 21, 2026Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
77 DeepSeek V3 0324 DeepSeek 55.0 64% confidence 64 percent, Medium not yet reported not yet reported 164Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 24, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
78 Claude Opus 4 Anthropic 54.8 96% confidence 96 percent, High not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 22, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
79 Grok 4.3 xAI 54.6 80% confidence 80 percent, High not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 17, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
80 Gemini 3 Flash Preview Google 54.4 85% confidence 85 percent, High $0.50Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$3.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 17, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
81 DeepSeek-V3 DeepSeek 54.1 75% confidence 75 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 26, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
82 GPT OSS 120B OpenAI 53.9 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
83 Claude Opus 4.1 Anthropic 53.5 80% confidence 80 percent, High not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
84 GLM-4.7-Flash Z.ai 53.3 69% confidence 69 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jan 19, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
85 Step 3.5 Flash stepfun 53.0 88% confidence 88 percent, High not yet reported not yet reported 256Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jan 29, 2026models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
86 Claude Sonnet 4.5 Anthropic 52.7 80% confidence 80 percent, High not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 29, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
87 Qwen3 30B A3B Alibaba / Qwen 52.6 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 28, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
88 GPT OSS 20B OpenAI 52.5 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
89 GPT-5.2 OpenAI 51.4 53% confidence 53 percent, Medium not yet reported not yet reported 400Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 11, 2025OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
90 Claude Sonnet 3.5 v2 Anthropic 51.3 75% confidence 75 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Oct 22, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
91 Llama 4 Scout 17B Instruct Meta 51.0 64% confidence 64 percent, Medium not yet reported not yet reported 10Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
92 GPT-5 OpenAI 50.9 53% confidence 53 percent, Medium not yet reported not yet reported 400Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 7, 2025OpenAI API changelogPublished source fact Retrieved Oct 8, 2026 · factual citation
Open source ↗
—
93 Qwen3 32B Alibaba / Qwen 50.4 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 1, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
94 Qwen3 235B-A22B Alibaba / Qwen 50.3 69% confidence 69 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 1, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
95 Claude Sonnet 3.7 Anthropic 49.8 80% confidence 80 percent, High not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Feb 19, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
96 GPT-4o (2024-05-13) OpenAI 49.4 75% confidence 75 percent, Medium not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 13, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
97 QwQ 32B Alibaba / Qwen 48.3 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
98 DeepSeek-R1 DeepSeek 48.0 85% confidence 85 percent, High not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jan 20, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
99 Llama-3.1-70B-Instruct Meta 45.8 75% confidence 75 percent, Medium not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 23, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
100 Gemini 2.5 Pro Google 45.5 85% confidence 85 percent, High $1.25Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
$10.00Google Gemini pricingOfficial paid Standard text rate, lowest short-context tier; current promotional price if dated; excludes free/Batch/Flex/audio Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jun 5, 2025Google Gemini release notesPublished source fact Retrieved Oct 8, 2026 · CC-BY-4.0 factual citation
Open source ↗
—
101 Claude Sonnet 4 Anthropic 45.2 96% confidence 96 percent, High not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 22, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
102 Gemma 3 12B IT Google 45.1 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 12, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
103 o3-mini OpenAI 43.5 75% confidence 75 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 20, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
104 Gemma 3 27B IT Google 43.1 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 12, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
105 Claude Haiku 3.5 Anthropic 41.7 75% confidence 75 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Oct 22, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
106 Claude Haiku 3 Anthropic 39.9 75% confidence 75 percent, Medium not yet reported not yet reported 200Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 13, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
107 Qwen2.5-Coder-32B-Instruct Alibaba / Qwen 39.6 75% confidence 75 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Nov 12, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
108 Gemma 3 4B IT Google 39.2 64% confidence 64 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Mar 12, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
109 GPT-4o (2024-08-06) OpenAI 38.8 88% confidence 88 percent, High not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Aug 6, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
110 Llama-3.1-8B-Instruct Meta 37.8 75% confidence 75 percent, Medium not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Jul 23, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
111 Mistral Large 2.1 mistral 37.4 88% confidence 88 percent, High not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Nov 18, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
112 Llama-3.3-70B-Instruct Meta 35.1 88% confidence 88 percent, High not yet reported not yet reported 128Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Dec 6, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
113 Mistral Medium 3 mistral 33.0 85% confidence 85 percent, High not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
May 7, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
—
114 Llama 4 Maverick 17B Instruct Meta 32.2 69% confidence 69 percent, Medium not yet reported not yet reported 1Mmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Apr 5, 2025models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open
115 Llama-3.2-1B Meta 32.2 75% confidence 75 percent, Medium not yet reported not yet reported 131Kmodels.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Sep 25, 2024models.devPublished source fact Retrieved Oct 8, 2026 · MIT
Open source ↗
Open

Ranked models appear first. Provisional models need more evidence and are available through the toggle (or the release catalog). Prices are USD per 1M tokens (official first-party API). Dotted values carry their source — hover or tap to see it. “not yet reported” means the source has not published that figure for this model.

Score against price, and over time

Best value models →
SI Score 40 60 80 100 $0.25 $0.5 $1 $2.5 $5 $10 USD / 1M tokens (log scale; 3:1 in/out) Claude Fable 5.1 — SI 80.2 · $20.00/1M blended · confidence 100% Claude Opus 5.5 — SI 78.4 · $8.00/1M blended · confidence 100% GPT-6 Astra — SI 77.6 · $20.00/1M blended · confidence 100% GPT-6.1 Sol — SI 73.0 · $4.00/1M blended · confidence 100% Gemini 3.7 Flash — SI 70.0 · $1.50/1M blended · confidence 100% Claude Sonnet 5.5 — SI 69.4 · $4.00/1M blended · confidence 93% Gemini 3.8 Flash — SI 68.8 · $1.50/1M blended · confidence 100% Gemini 3.5 Flash — SI 68.3 · $3.38/1M blended · confidence 100% Gemini 3.1 Pro Preview — SI 68.1 · $4.50/1M blended · confidence 100% DeepSeek V4.1 Flash — SI 66.2 · $0.26/1M blended · confidence 69% Gemini 3.6 Flash — SI 65.6 · $1.50/1M blended · confidence 100% Grok 4.7 — SI 65.5 · $3.00/1M blended · confidence 100% Claude Haiku 5.5 — SI 64.8 · $0.20/1M blended · confidence 100% DeepSeek V4 Pro 0813 — SI 63.0 · $0.99/1M blended · confidence 69% GPT-6 Luna — SI 61.6 · $0.20/1M blended · confidence 100% Gemini 3.5 Flash Lite — SI 55.2 · $0.85/1M blended · confidence 100% Gemini 3 Flash Preview — SI 54.4 · $1.13/1M blended · confidence 85% Gemini 2.5 Pro — SI 45.5 · $3.44/1M blended · confidence 85%
Closed weights Open weights Dot opacity = confidence in the score
SI Score vs blended API price (log scale). Upper-left is the sweet spot: more score per dollar.
SI Score 40 60 80 100 Mar ’24 Oct ’24 May ’25 Dec ’25 Jul ’26 Release date Claude Fable 5.1 — SI 80.2 · released 2026-09-01 · confidence 100% Claude Opus 5.5 — SI 78.4 · released 2026-09-22 · confidence 100% GPT-6 Astra — SI 77.6 · released 2026-09-03 · confidence 100% Claude Fable 5 — SI 76.8 · released 2026-06-09 · confidence 100% Claude Opus 5 — SI 74.9 · released 2026-07-24 · confidence 100% GPT-6.1 Sol — SI 73.0 · released 2026-09-29 · confidence 100% GPT-5.6 Sol — SI 72.1 · released 2026-07-09 · confidence 100% Kimi K3 — SI 71.0 · released 2026-07-16 · confidence 100% Claude Opus 4.6 — SI 70.5 · released 2026-02-05 · confidence 100% Muse Spark 1.3 — SI 70.1 · released 2026-09-02 · confidence 93% Claude Opus 4.7 — SI 70.1 · released 2026-04-16 · confidence 100% Gemini 3.7 Flash — SI 70.0 · released 2026-08-13 · confidence 100% GPT-5.5 — SI 69.5 · released 2026-04-24 · confidence 100% Claude Sonnet 5.5 — SI 69.4 · released 2026-09-28 · confidence 93% GPT-6 Sol — SI 68.9 · released 2026-09-22 · confidence 100% Gemini 3.8 Flash — SI 68.8 · released 2026-09-02 · confidence 100% Gemini 3.5 Flash — SI 68.3 · released 2026-05-19 · confidence 100% Claude Sonnet 4.6 — SI 68.2 · released 2026-02-17 · confidence 100% GPT-5.6 Terra — SI 68.2 · released 2026-07-09 · confidence 100% Gemini 3.1 Pro Preview — SI 68.1 · released 2026-02-19 · confidence 100% MiniMax-M3 — SI 68.0 · released 2026-06-01 · confidence 100% GLM-5.3 — SI 67.7 · released 2026-08-14 · confidence 93% Claude Sonnet 5 — SI 67.7 · released 2026-06-30 · confidence 93% Claude Opus 4.8 — SI 67.5 · released 2026-05-28 · confidence 100% Grok 4.6 — SI 66.6 · released 2026-08-12 · confidence 100% Qwen3.6 Max Preview — SI 66.4 · released 2026-04-20 · confidence 64% DeepSeek V4.1 Flash — SI 66.2 · released 2026-09-10 · confidence 69% GLM-5.3-Flash — SI 66.2 · released 2026-08-26 · confidence 100% Kimi K2 Thinking Turbo — SI 66.0 · released 2025-11-06 · confidence 64% GLM-5.2 — SI 65.7 · released 2026-06-13 · confidence 100% Gemini 3.6 Flash — SI 65.6 · released 2026-07-21 · confidence 100% Grok 4.7 — SI 65.5 · released 2026-09-21 · confidence 100% Grok 4.5 — SI 64.8 · released 2026-07-08 · confidence 100% Claude Haiku 5.5 — SI 64.8 · released 2026-10-07 · confidence 100% Claude Opus 4.5 — SI 64.6 · released 2025-11-01 · confidence 100% Kimi K2.6 — SI 64.0 · released 2026-04-21 · confidence 85% Qwen3.5 397B-A17B — SI 64.0 · released 2026-02-15 · confidence 69% GPT-5.5 Pro — SI 63.5 · released 2026-04-24 · confidence 53% GPT-5.6 Luna — SI 63.4 · released 2026-07-09 · confidence 100% Gemma 4 26B A4B IT — SI 63.4 · released 2026-04-02 · confidence 64% GLM-5.1 — SI 63.3 · released 2026-04-07 · confidence 64% DeepSeek V4 Pro — SI 63.1 · released 2026-04-24 · confidence 85% Gemma 4 31B IT — SI 63.0 · released 2026-04-02 · confidence 64% DeepSeek V4 Pro 0813 — SI 63.0 · released 2026-08-12 · confidence 69% Qwen3.7 Plus — SI 62.7 · released 2026-06-02 · confidence 64% Qwen3.5 35B-A3B — SI 62.7 · released 2026-02-23 · confidence 64% MiMo-V2.5-Pro — SI 62.6 · released 2026-04-22 · confidence 88% Muse Spark 1.2 — SI 62.3 · released 2026-08-05 · confidence 53% Qwen3.8 Max — SI 62.0 · released 2026-08-03 · confidence 53% GPT-5.4 — SI 61.9 · released 2026-03-05 · confidence 69% Qwen3.8 27B — SI 61.9 · released 2026-08-14 · confidence 69% GPT-6 Luna — SI 61.6 · released 2026-09-22 · confidence 100% Qwen3.6 Plus — SI 61.6 · released 2026-04-02 · confidence 80% GLM-4.7 — SI 61.3 · released 2025-12-22 · confidence 69% Muse Spark 1.1 — SI 61.2 · released 2026-04-08 · confidence 53% DeepSeek V4 Flash 0731 — SI 61.0 · released 2026-07-31 · confidence 69% Hy3 — SI 60.7 · released 2026-07-06 · confidence 88% Qwen3.5 Flash — SI 59.5 · released 2026-02-23 · confidence 64% GLM-5 — SI 58.9 · released 2026-02-12 · confidence 85% Kimi K2.5 — SI 58.8 · released 2026-01 · confidence 100% Qwen3.7 Max — SI 58.6 · released 2026-05-21 · confidence 53% GPT-5.5 Instant — SI 58.3 · released 2026-05-05 · confidence 64% Claude Haiku 4.5 — SI 58.1 · released 2025-10-15 · confidence 64% MiniMax-M2.7 — SI 57.9 · released 2026-03-18 · confidence 88% Grok 4.20 (Reasoning) — SI 57.7 · released 2026-03-09 · confidence 80% MiMo-V2.6-Pro — SI 57.4 · released 2026-09-22 · confidence 88% Gemini 3 Pro Preview — SI 56.8 · released 2025-11-18 · confidence 53% MiniMax-M2.5 — SI 56.7 · released 2026-02-12 · confidence 88% GPT-5.4 nano — SI 56.6 · released 2026-03-17 · confidence 69% MiMo-V2.5 — SI 56.4 · released 2026-04-22 · confidence 88% DeepSeek V4 Flash — SI 56.2 · released 2026-04-24 · confidence 53% GPT-5.4 mini — SI 55.8 · released 2026-03-17 · confidence 69% MiMo-V2.6-Flash — SI 55.8 · released 2026-09-22 · confidence 88% MiniMax-M2 — SI 55.7 · released 2025-10-27 · confidence 88% Qwen3.6 27B — SI 55.6 · released 2026-04-22 · confidence 53% Gemini 3.5 Flash Lite — SI 55.2 · released 2026-07-21 · confidence 100% DeepSeek V3 0324 — SI 55.0 · released 2025-03-24 · confidence 64% Claude Opus 4 — SI 54.8 · released 2025-05-22 · confidence 96% Grok 4.3 — SI 54.6 · released 2026-04-17 · confidence 80% Gemini 3 Flash Preview — SI 54.4 · released 2025-12-17 · confidence 85% DeepSeek-V3 — SI 54.1 · released 2024-12-26 · confidence 75% GPT OSS 120B — SI 53.9 · released 2025-08-05 · confidence 64% Claude Opus 4.1 — SI 53.5 · released 2025-08-05 · confidence 80% GLM-4.7-Flash — SI 53.3 · released 2026-01-19 · confidence 69% Step 3.5 Flash — SI 53.0 · released 2026-01-29 · confidence 88% Claude Sonnet 4.5 — SI 52.7 · released 2025-09-29 · confidence 80% Qwen3 30B A3B — SI 52.6 · released 2025-04-28 · confidence 64% GPT OSS 20B — SI 52.5 · released 2025-08-05 · confidence 64% GPT-5.2 — SI 51.4 · released 2025-12-11 · confidence 53% Claude Sonnet 3.5 v2 — SI 51.3 · released 2024-10-22 · confidence 75% Llama 4 Scout 17B Instruct — SI 51.0 · released 2025-04-05 · confidence 64% GPT-5 — SI 50.9 · released 2025-08-07 · confidence 53% Qwen3 32B — SI 50.4 · released 2025-04 · confidence 64% Qwen3 235B-A22B — SI 50.3 · released 2025-04 · confidence 69% Claude Sonnet 3.7 — SI 49.8 · released 2025-02-19 · confidence 80% GPT-4o (2024-05-13) — SI 49.4 · released 2024-05-13 · confidence 75% QwQ 32B — SI 48.3 · released 2025-03-05 · confidence 64% DeepSeek-R1 — SI 48.0 · released 2025-01-20 · confidence 85% Llama-3.1-70B-Instruct — SI 45.8 · released 2024-07-23 · confidence 75% Gemini 2.5 Pro — SI 45.5 · released 2025-06-05 · confidence 85% Claude Sonnet 4 — SI 45.2 · released 2025-05-22 · confidence 96% Gemma 3 12B IT — SI 45.1 · released 2025-03-12 · confidence 64% o3-mini — SI 43.5 · released 2024-12-20 · confidence 75% Gemma 3 27B IT — SI 43.1 · released 2025-03-12 · confidence 64% Claude Haiku 3.5 — SI 41.7 · released 2024-10-22 · confidence 75% Claude Haiku 3 — SI 39.9 · released 2024-03-13 · confidence 75% Qwen2.5-Coder-32B-Instruct — SI 39.6 · released 2024-11-12 · confidence 75% Gemma 3 4B IT — SI 39.2 · released 2025-03-12 · confidence 64% GPT-4o (2024-08-06) — SI 38.8 · released 2024-08-06 · confidence 88% Llama-3.1-8B-Instruct — SI 37.8 · released 2024-07-23 · confidence 75% Mistral Large 2.1 — SI 37.4 · released 2024-11-18 · confidence 88% Llama-3.3-70B-Instruct — SI 35.1 · released 2024-12-06 · confidence 88% Mistral Medium 3 — SI 33.0 · released 2025-05-07 · confidence 85% Llama 4 Maverick 17B Instruct — SI 32.2 · released 2025-04-05 · confidence 69% Llama-3.2-1B — SI 32.2 · released 2024-09-25 · confidence 75%
Closed weights Open weights Dot opacity = confidence in the score
Ranked models by release date and current SI Score. This is a release comparison, not historical score tracking.

Latest releases

All releases →

Understand the numbers