arc agi v1 public eval Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Fable 5.1 Anthropic 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
99.0 100% confidence 100 percent, Full
2 Claude Opus 5 Anthropic 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
99.0 100% confidence 100 percent, Full
3 Claude Opus 5 Anthropic 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
99.0 100% confidence 100 percent, Full
4 GPT-5.6 Sol OpenAI 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
99.0 100% confidence 100 percent, Full
5 GPT-6 Astra OpenAI 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
99.0 100% confidence 100 percent, Full
6 Claude Fable 5.1 Anthropic 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.8 100% confidence 100 percent, Full
7 GPT-5.6 Sol OpenAI 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.8 100% confidence 100 percent, Full
8 GPT-6 Astra OpenAI 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.8 100% confidence 100 percent, Full
9 Claude Opus 5.5 Anthropic 98.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.6 100% confidence 100 percent, Full
10 GPT-6 Astra OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
11 GPT-6.1 Sol OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
12 GPT-6.1 Sol OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
13 Claude Opus 5.5 Anthropic 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
14 Claude Opus 5.5 Anthropic 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
15 Claude Opus 5.5 Anthropic 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
16 GPT-5.4 Pro OpenAI 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 69% confidence 69 percent, Medium
17 GPT-5.6 Terra OpenAI 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
18 GPT-6 Astra OpenAI 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
19 GPT-6.1 Sol OpenAI 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.3 100% confidence 100 percent, Full
20 Claude Fable 5 Anthropic 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 100% confidence 100 percent, Full
21 Claude Fable 5 Anthropic 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 100% confidence 100 percent, Full
22 DeepSeek V4.1 Flash DeepSeek 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 69% confidence 69 percent, Medium
23 Gemini 3.8 Flash Google 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 100% confidence 100 percent, Full
24 GPT-5.5 Pro OpenAI 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 53% confidence 53 percent, Medium
25 GPT-5.5 Pro OpenAI 98.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 53% confidence 53 percent, Medium
26 Claude Fable 5 Anthropic 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.9 100% confidence 100 percent, Full
27 GPT-6.1 Sol OpenAI 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.9 100% confidence 100 percent, Full
28 Gemini 3.8 Flash Google 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.8 100% confidence 100 percent, Full
29 GPT-6 Astra OpenAI 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.8 100% confidence 100 percent, Full
30 GPT-6 Sol OpenAI 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.8 100% confidence 100 percent, Full
31 Claude Fable 5 Anthropic 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.6 100% confidence 100 percent, Full
32 Claude Fable 5.1 Anthropic 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.6 100% confidence 100 percent, Full
33 GPT-5.2 Pro OpenAI 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.6 48% confidence 48 percent, Low
34 GPT-5.5 OpenAI 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
35 GPT-5.5 OpenAI 97.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.4 100% confidence 100 percent, Full
36 GPT-5.6 Sol OpenAI 97.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.3 100% confidence 100 percent, Full
37 Gemini 3.1 Pro Preview Google 97.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.2 100% confidence 100 percent, Full
38 Claude Opus 4.7 Anthropic 97%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.0 100% confidence 100 percent, Full
39 Claude Opus 4.6 Anthropic 96.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.8 100% confidence 100 percent, Full
40 Claude Opus 4.7 Anthropic 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.6 100% confidence 100 percent, Full
41 Claude Fable 5.1 Anthropic 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.6 100% confidence 100 percent, Full
42 Gemini 3.7 Flash Google 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.6 100% confidence 100 percent, Full
43 Claude Fable 5 Anthropic 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
44 GPT-5.5 OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
45 GPT-6 Sol OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
46 GPT-5.4 OpenAI 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.4 100% confidence 100 percent, Full
47 GPT-6.1 Sol OpenAI 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.4 100% confidence 100 percent, Full
48 Claude Opus 4.6 Anthropic 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.3 100% confidence 100 percent, Full
49 DeepSeek V4.1 Flash DeepSeek 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.3 69% confidence 69 percent, Medium
50 Grok 4.6 xAI 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.3 100% confidence 100 percent, Full
51 Grok 4.6 xAI 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.3 100% confidence 100 percent, Full
52 Claude Opus 4.7 Anthropic 96.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.1 100% confidence 100 percent, Full
53 Gemini 3.5 Flash Google 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.0 100% confidence 100 percent, Full
54 Gemini 3.6 Flash Google 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.0 100% confidence 100 percent, Full
55 DeepSeek V4.1 Flash DeepSeek 95.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.9 69% confidence 69 percent, Medium
56 Claude Sonnet 4.6 Anthropic 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.8 100% confidence 100 percent, Full
57 GPT-5.6 Terra OpenAI 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.8 100% confidence 100 percent, Full
58 Grok 4.7 xAI 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.8 100% confidence 100 percent, Full
59 GPT-5.6 Terra OpenAI 95.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.6 100% confidence 100 percent, Full
60 GPT-5.4 OpenAI 95.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.6 100% confidence 100 percent, Full
61 Claude Fable 5.1 Anthropic 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 100% confidence 100 percent, Full
62 DeepSeek V4 Pro 0813 DeepSeek 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 69% confidence 69 percent, Medium
63 Grok 4.20 (Reasoning) xAI 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 80% confidence 80 percent, High
64 Claude Sonnet 4.6 Anthropic 95.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.3 100% confidence 100 percent, Full
65 Gemini 3.7 Flash Google 95.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.3 100% confidence 100 percent, Full
66 Grok 4.7 xAI 95.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.1 100% confidence 100 percent, Full
67 GPT-5.2 OpenAI 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.0 100% confidence 100 percent, Full
68 Grok 4.7 xAI 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.0 100% confidence 100 percent, Full
69 Kimi K3 Moonshot AI 94.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.9 100% confidence 100 percent, Full
70 Claude Opus 4.6 Anthropic 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.8 100% confidence 100 percent, Full
71 DeepSeek V4 Flash 0731 DeepSeek 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.8 69% confidence 69 percent, Medium
72 Grok 4.5 xAI 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.8 100% confidence 100 percent, Full
73 Grok 4.6 xAI 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.8 100% confidence 100 percent, Full
74 GPT-5.2 Pro OpenAI 94.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.6 48% confidence 48 percent, Low
75 Claude Opus 4.7 Anthropic 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 100% confidence 100 percent, Full
76 Gemini 3.8 Flash Google 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 100% confidence 100 percent, Full
77 GPT-5.6 Sol OpenAI 93.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.9 100% confidence 100 percent, Full
78 DeepSeek V4 Pro 0813 DeepSeek 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.3 69% confidence 69 percent, Medium
79 DeepSeek V4 Flash 0731 DeepSeek 93%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.0 69% confidence 69 percent, Medium
80 Kimi K3 Moonshot AI 92.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.9 100% confidence 100 percent, Full
81 DeepSeek V4 Pro 0813 DeepSeek 92.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.8 69% confidence 69 percent, Medium
82 Gemini 3.6 Flash Google 92.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.8 100% confidence 100 percent, Full
83 DeepSeek V4 Flash 0731 DeepSeek 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 69% confidence 69 percent, Medium
84 GPT-6 Luna OpenAI 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
85 Grok 4.5 xAI 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
86 GLM-5.3-Flash Z.ai 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
87 Claude Opus 5.5 Anthropic 92.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.3 100% confidence 100 percent, Full
88 GPT-5.4 OpenAI 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.0 100% confidence 100 percent, Full
89 GPT-6 Sol OpenAI 91.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.4 100% confidence 100 percent, Full
90 Gemini 3.7 Flash Google 91.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.3 100% confidence 100 percent, Full
91 GPT-5.6 Luna OpenAI 90.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.8 100% confidence 100 percent, Full
92 Inkling thinkingmachines 90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.6 100% confidence 100 percent, Full
93 GPT-5.2 OpenAI 90.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.3 100% confidence 100 percent, Full
94 GPT-5.2 Pro OpenAI 90.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.3 48% confidence 48 percent, Low
95 GPT-5.6 Luna OpenAI 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
96 Claude Opus 4.6 Anthropic 89.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.6 100% confidence 100 percent, Full
97 Inkling Small thinkingmachines 89.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.4 100% confidence 100 percent, Full
98 GPT-6 Sol OpenAI 88.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.9 100% confidence 100 percent, Full
99 Gemini 3 Flash Preview Google 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.3 93% confidence 93 percent, High
100 Qwen3.8 27B Alibaba / Qwen 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.5 69% confidence 69 percent, Medium
101 Inkling Small thinkingmachines 86.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.6 100% confidence 100 percent, Full
102 Claude Opus 4.5 Anthropic 86.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.6 100% confidence 100 percent, Full
103 Grok 4.5 xAI 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.0 100% confidence 100 percent, Full
104 GPT-6 Luna OpenAI 85.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.8 100% confidence 100 percent, Full
105 GPT-5.6 Terra OpenAI 84.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.8 100% confidence 100 percent, Full
106 GPT-5.6 Sol OpenAI 84.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.5 100% confidence 100 percent, Full
107 Grok 4.6 xAI 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.0 100% confidence 100 percent, Full
108 GPT-5.5 OpenAI 82.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
82.3 100% confidence 100 percent, Full
109 Grok 4.7 xAI 82.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
82.3 100% confidence 100 percent, Full
110 Claude Opus 4.5 Anthropic 81.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
81.6 100% confidence 100 percent, Full
111 Gemini 3.6 Flash Google 81.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
81.4 100% confidence 100 percent, Full
112 GPT-5.2 OpenAI 80.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
80.6 100% confidence 100 percent, Full
113 GLM-5.2 Z.ai 80.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
80.4 100% confidence 100 percent, Full
114 GPT-5.4 OpenAI 80%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
80.0 100% confidence 100 percent, Full
115 GPT-6 Luna OpenAI 79.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
79.8 100% confidence 100 percent, Full
116 GPT-5.6 Luna OpenAI 79.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
79.3 100% confidence 100 percent, Full
117 GPT-6 Sol OpenAI 79.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
79.3 100% confidence 100 percent, Full
118 GLM-5.3-Flash Z.ai 78.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
78.6 100% confidence 100 percent, Full
119 Inkling Small thinkingmachines 77.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.8 100% confidence 100 percent, Full
120 GPT-5.1 OpenAI 77.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.1 85% confidence 85 percent, High
121 GPT-5 Pro OpenAI 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.0 64% confidence 64 percent, Medium
122 Qwen3.8 27B Alibaba / Qwen 76.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.3 69% confidence 69 percent, Medium
123 Qwen3.8 27B Alibaba / Qwen 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
75.8 69% confidence 69 percent, Medium
124 Kimi K3 Moonshot AI 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
75.8 100% confidence 100 percent, Full
125 GPT-5.4 mini OpenAI 75.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
75.1 100% confidence 100 percent, Full
126 Claude Sonnet 4.5 (latest) Anthropic 73.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
73.8 43% confidence 43 percent, Low
127 Kimi K2.5 Moonshot AI 73.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
73.1 100% confidence 100 percent, Full
128 GPT-5.6 Terra OpenAI 70.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.8 100% confidence 100 percent, Full
129 GPT-6 Luna OpenAI 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.1 100% confidence 100 percent, Full
130 Claude Opus 4.5 Anthropic 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.1 100% confidence 100 percent, Full
131 GPT-5.1 OpenAI 68.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.9 85% confidence 85 percent, High
132 o4-mini OpenAI 68.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.0 87% confidence 87 percent, High
133 Gemini 3 Flash Preview Google 67.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.9 93% confidence 93 percent, High
134 Gemini 3.5 Flash Lite Google 66.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
66.4 100% confidence 100 percent, Full
135 GPT-5.4 mini OpenAI 66.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
66.3 100% confidence 100 percent, Full
136 Inkling Small thinkingmachines 66.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
66.1 100% confidence 100 percent, Full
137 GPT-5.2 OpenAI 65.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.9 100% confidence 100 percent, Full
138 GPT-5 OpenAI 65.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.9 100% confidence 100 percent, Full
139 GPT-5.6 Luna OpenAI 64.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
64.4 100% confidence 100 percent, Full
140 o3 OpenAI 64.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
64.3 87% confidence 87 percent, High
141 Claude Sonnet 4.5 (latest) Anthropic 63.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.6 43% confidence 43 percent, Low
142 GPT-5 OpenAI 63.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.4 100% confidence 100 percent, Full
143 o3-pro OpenAI 63.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.3 22% confidence 22 percent, Low
144 Claude Haiku 4.5 (latest) Anthropic 62.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
62.9 33% confidence 33 percent, Low
145 DeepSeek V3.2 DeepSeek 61.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.6 56% confidence 56 percent, Medium
146 GPT-5 Mini OpenAI 61.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.5 99% confidence 99 percent, High
147 MiniMax-M2.5 minimax 59.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.1 100% confidence 100 percent, Full
148 GLM-5 Z.ai 58.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.6 90% confidence 90 percent, High
149 o3-pro OpenAI 58.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.1 22% confidence 22 percent, Low
150 GLM-5.3-Flash Z.ai 57.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.8 100% confidence 100 percent, Full
151 Claude Sonnet 4 Anthropic 56.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.8 100% confidence 100 percent, Full
152 o3 OpenAI 56.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.7 87% confidence 87 percent, High
153 Gemini 2.5 Pro Google 56.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.4 90% confidence 90 percent, High
154 Gemini 2.5 Pro Google 55.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
55.9 90% confidence 90 percent, High
155 GPT-5.4 mini OpenAI 55.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
55.4 100% confidence 100 percent, Full
156 Claude Opus 4 Anthropic 54.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
54.3 100% confidence 100 percent, Full
157 Claude Sonnet 4.5 (latest) Anthropic 53.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
53.5 43% confidence 43 percent, Low
158 Claude Opus 4.5 Anthropic 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.6 100% confidence 100 percent, Full
159 GPT-5.4 nano OpenAI 51.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
51.6 100% confidence 100 percent, Full
160 Claude Haiku 4.5 (latest) Anthropic 51.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
51.4 33% confidence 33 percent, Low
161 o3-pro OpenAI 50.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
50.9 22% confidence 22 percent, Low
162 o4-mini OpenAI 50.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
50.2 87% confidence 87 percent, High
163 Claude Sonnet 4 Anthropic 48.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
48.6 100% confidence 100 percent, Full
164 GPT-5 OpenAI 48.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
48.4 100% confidence 100 percent, Full
165 GPT-5.4 nano OpenAI 47.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.9 100% confidence 100 percent, Full
166 o3 OpenAI 47.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.6 87% confidence 87 percent, High
167 GPT-5.6 Luna OpenAI 47.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.4 100% confidence 100 percent, Full
168 Gemini 3.5 Flash Lite Google 47%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.0 100% confidence 100 percent, Full
169 o3-mini OpenAI 46.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.6 96% confidence 96 percent, High
170 GPT-6 Luna OpenAI 46.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.5 100% confidence 100 percent, Full
171 GPT-5 Mini OpenAI 46.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.3 99% confidence 99 percent, High
172 Claude Opus 4 Anthropic 45.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
45.6 100% confidence 100 percent, Full
173 Claude Haiku 4.5 (latest) Anthropic 45%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
45.0 33% confidence 33 percent, Low
174 Gemini 2.5 Pro Google 44.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.2 90% confidence 90 percent, High
175 GPT-5.1 OpenAI 44%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.0 85% confidence 85 percent, High
176 GPT-5.4 nano OpenAI 43.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
43.4 100% confidence 100 percent, Full
177 Claude Opus 4 Anthropic 43.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
43.3 100% confidence 100 percent, Full
178 Gemini 3 Flash Preview Google 38.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
38.2 93% confidence 93 percent, High
179 Claude Sonnet 4.5 (latest) Anthropic 36.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
36.6 43% confidence 43 percent, Low
180 Claude Opus 4 Anthropic 35.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
35.5 100% confidence 100 percent, Full
181 Claude Sonnet 4.5 (latest) Anthropic 35.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
35.4 43% confidence 43 percent, Low
182 Claude Sonnet 4 Anthropic 33%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.0 100% confidence 100 percent, Full
183 GPT-5.4 mini OpenAI 31.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.8 100% confidence 100 percent, Full
184 Claude Sonnet 4 Anthropic 31.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.3 100% confidence 100 percent, Full
185 o3-mini OpenAI 30.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
30.6 96% confidence 96 percent, High
186 GPT-5 Nano OpenAI 29.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.7 85% confidence 85 percent, High
187 o4-mini OpenAI 27.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.6 87% confidence 87 percent, High
188 Claude Haiku 4.5 (latest) Anthropic 27.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.1 33% confidence 33 percent, Low
189 DeepSeek-R1 DeepSeek 27.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.0 88% confidence 88 percent, High
190 Claude Haiku 4.5 (latest) Anthropic 26.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
26.6 33% confidence 33 percent, Low
191 Gemini 3.5 Flash Lite Google 24.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
24.6 100% confidence 100 percent, Full
192 GPT-5.4 nano OpenAI 24.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
24.6 100% confidence 100 percent, Full
193 GPT-5 Mini OpenAI 24.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
24.4 99% confidence 99 percent, High
194 GPT-5 Nano OpenAI 20.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
20.8 85% confidence 85 percent, High
195 Gemini 2.5 Pro Google 17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.5 90% confidence 90 percent, High
196 o3-mini OpenAI 17.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.4 96% confidence 96 percent, High
197 Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen 17%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.0 48% confidence 48 percent, Low
198 GPT-5.2 OpenAI 16.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
16.5 100% confidence 100 percent, Full
199 GPT-5.1 OpenAI 12.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
12.4 85% confidence 85 percent, High
200 GPT-5 Nano OpenAI 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
11.8 85% confidence 85 percent, High
201 GPT-4.1 OpenAI 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
11.8 100% confidence 100 percent, Full
202 Magistral Medium (latest) mistral 8.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.9 16% confidence 16 percent, Low
203 Magistral Small mistral 8.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.6 48% confidence 48 percent, Low
204 Magistral Medium (latest) mistral 8.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.0 16% confidence 16 percent, Low
205 GPT-4.1 mini OpenAI 7.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.2 87% confidence 87 percent, High
206 Llama 4 Maverick 17B Instruct Meta 7.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.1 78% confidence 78 percent, Medium
207 GPT-4.1 nano OpenAI 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.8 82% confidence 82 percent, High

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.