arc agi v1 semi private Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 2 | Claude Fable 5 Anthropic | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 4 | Gemini 3.8 Flash Google | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 5 | GPT-6 Astra OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 6 | GPT-6 Astra OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 7 | GPT-6.1 Sol OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 8 | GPT-6.1 Sol OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 9 | Gemini 3.1 Pro Preview Google | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 100% confidence 100 percent, Full |
| 10 | Claude Fable 5.1 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 11 | Claude Opus 5 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 12 | Claude Opus 5 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 13 | Claude Opus 5.5 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 14 | Claude Opus 5.5 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 15 | Claude Opus 5.5 Anthropic | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 16 | Gemini 3.8 Flash Google | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 17 | GPT-5.6 Sol OpenAI | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 18 | GPT-6 Astra OpenAI | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 19 | GPT-6 Astra OpenAI | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 20 | GPT-5.6 Sol OpenAI | 97%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.0 | 100% confidence 100 percent, Full |
| 21 | Claude Fable 5.1 Anthropic | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 22 | GPT-5.5 Pro OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 53% confidence 53 percent, Medium |
| 23 | GPT-5.6 Sol OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 24 | GPT-5.6 Terra OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 25 | GPT-6 Astra OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 26 | GPT-6.1 Sol OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 27 | Claude Fable 5.1 Anthropic | 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.0 | 100% confidence 100 percent, Full |
| 28 | Claude Fable 5 Anthropic | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 29 | Gemini 3.7 Flash Google | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 30 | GPT-6 Sol OpenAI | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 31 | GPT-6.1 Sol OpenAI | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 32 | GPT-5.5 OpenAI | 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.0 | 100% confidence 100 percent, Full |
| 33 | GPT-5.5 Pro OpenAI | 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.0 | 53% confidence 53 percent, Medium |
| 34 | Claude Fable 5.1 Anthropic | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 100% confidence 100 percent, Full |
| 35 | DeepSeek V4.1 Flash DeepSeek | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 69% confidence 69 percent, Medium |
| 36 | Kimi K3 Moonshot AI | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 100% confidence 100 percent, Full |
| 37 | GPT-5.4 Pro OpenAI | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 69% confidence 69 percent, Medium |
| 38 | GPT-5.5 OpenAI | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 100% confidence 100 percent, Full |
| 39 | Claude Opus 4.6 Anthropic | 94%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.0 | 100% confidence 100 percent, Full |
| 40 | GPT-5.6 Terra OpenAI | 94%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.0 | 100% confidence 100 percent, Full |
| 41 | GPT-5.4 OpenAI | 93.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.7 | 100% confidence 100 percent, Full |
| 42 | Claude Opus 4.7 Anthropic | 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.5 | 100% confidence 100 percent, Full |
| 43 | GPT-6.1 Sol OpenAI | 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.5 | 100% confidence 100 percent, Full |
| 44 | Claude Opus 4.6 Anthropic | 93%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.0 | 100% confidence 100 percent, Full |
| 45 | GPT-5.4 OpenAI | 92.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.7 | 100% confidence 100 percent, Full |
| 46 | GPT-6 Sol OpenAI | 92.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.7 | 100% confidence 100 percent, Full |
| 47 | Claude Fable 5 Anthropic | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 48 | Claude Opus 4.8 Anthropic | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-opus-4-8-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 49 | Gemini 3.5 Flash Google | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 50 | GPT-5.6 Sol OpenAI | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 51 | GPT-5.5 OpenAI | 92.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.2 | 100% confidence 100 percent, Full |
| 52 | Claude Opus 4.6 Anthropic | 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.0 | 100% confidence 100 percent, Full |
| 53 | Claude Opus 4.7 Anthropic | 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.0 | 100% confidence 100 percent, Full |
| 54 | Claude Opus 4.8 Anthropic | 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.0 | 100% confidence 100 percent, Full |
| 55 | GPT-5.6 Terra OpenAI | 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.0 | 100% confidence 100 percent, Full |
| 56 | Claude Opus 4.8 Anthropic | 91.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' output effort. [variant] anthropic-opus-4-8-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.5 | 100% confidence 100 percent, Full |
| 57 | Gemini 3.6 Flash Google | 91.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.2 | 100% confidence 100 percent, Full |
| 58 | Gemini 3.7 Flash Google | 91.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.2 | 100% confidence 100 percent, Full |
| 59 | Claude Opus 4.7 Anthropic | 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.0 | 100% confidence 100 percent, Full |
| 60 | Claude Opus 4.7 Anthropic | 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.0 | 100% confidence 100 percent, Full |
| 61 | GPT-6 Sol OpenAI | 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.0 | 100% confidence 100 percent, Full |
| 62 | GLM-5.3-Flash Z.ai | 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.0 | 100% confidence 100 percent, Full |
| 63 | Claude Fable 5 Anthropic | 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.5 | 100% confidence 100 percent, Full |
| 64 | DeepSeek V4 Pro 0813 DeepSeek | 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.5 | 69% confidence 69 percent, Medium |
| 65 | DeepSeek V4.1 Flash DeepSeek | 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.5 | 69% confidence 69 percent, Medium |
| 66 | Gemini 3.8 Flash Google | 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.5 | 100% confidence 100 percent, Full |
| 67 | GPT-5.2 Pro OpenAI | 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.5 | 48% confidence 48 percent, Low |
| 68 | Grok 4.7 xAI | 90.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.2 | 100% confidence 100 percent, Full |
| 69 | Claude Fable 5.1 Anthropic | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 70 | DeepSeek V4 Pro 0813 DeepSeek | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 69% confidence 69 percent, Medium |
| 71 | Grok 4.7 xAI | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 72 | Grok 4.20 (Reasoning) xAI | 89.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.5 | 80% confidence 80 percent, High |
| 73 | Grok 4.7 xAI | 89.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.5 | 100% confidence 100 percent, Full |
| 74 | DeepSeek V4 Flash 0731 DeepSeek | 89%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.0 | 69% confidence 69 percent, Medium |
| 75 | Claude Opus 5.5 Anthropic | 88.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.5 | 100% confidence 100 percent, Full |
| 76 | DeepSeek V4.1 Flash DeepSeek | 88.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.5 | 69% confidence 69 percent, Medium |
| 77 | Claude Opus 4.8 Anthropic | 88%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' output effort. [variant] anthropic-opus-4-8-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.0 | 100% confidence 100 percent, Full |
| 78 | GPT-5.6 Luna OpenAI | 88%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.0 | 100% confidence 100 percent, Full |
| 79 | GPT-5.6 Luna OpenAI | 87.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.7 | 100% confidence 100 percent, Full |
| 80 | Qwen3.8 27B Alibaba / Qwen | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 69% confidence 69 percent, Medium |
| 81 | Grok 4.6 xAI | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 82 | Grok 4.5 xAI | 87.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.2 | 100% confidence 100 percent, Full |
| 83 | DeepSeek V4 Pro 0813 DeepSeek | 87.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.2 | 69% confidence 69 percent, Medium |
| 84 | DeepSeek V4 Flash 0731 DeepSeek | 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.0 | 69% confidence 69 percent, Medium |
| 85 | Grok 4.6 xAI | 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 86 | Grok 4.6 xAI | 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 87 | GPT-6 Luna OpenAI | 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.7 | 100% confidence 100 percent, Full |
| 88 | Kimi K3 Moonshot AI | 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.7 | 100% confidence 100 percent, Full |
| 89 | Claude Sonnet 4.6 Anthropic | 86.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.5 | 100% confidence 100 percent, Full |
| 90 | GPT-5.2 OpenAI | 86.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.2 | 100% confidence 100 percent, Full |
| 91 | GPT-5.4 OpenAI | 86.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.2 | 100% confidence 100 percent, Full |
| 92 | Claude Opus 4.6 Anthropic | 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 93 | Claude Sonnet 4.6 Anthropic | 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 94 | GPT-5.2 Pro OpenAI | 85.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.7 | 48% confidence 48 percent, Low |
| 95 | Grok 4.5 xAI | 85.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.7 | 100% confidence 100 percent, Full |
| 96 | Gemini 3.7 Flash Google | 85.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.2 | 100% confidence 100 percent, Full |
| 97 | Gemini 3 Flash Preview Google | 84.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.7 | 93% confidence 93 percent, High |
| 98 | DeepSeek V4 Flash 0731 DeepSeek | 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.0 | 69% confidence 69 percent, Medium |
| 99 | Inkling Small thinkingmachines | 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.0 | 100% confidence 100 percent, Full |
| 100 | GPT-6 Sol OpenAI | 83.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 83.7 | 100% confidence 100 percent, Full |
| 101 | Gemini 3.6 Flash Google | 83.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 83.2 | 100% confidence 100 percent, Full |
| 102 | GPT-5.2 Pro OpenAI | 81.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.2 | 48% confidence 48 percent, Low |
| 103 | Claude Opus 4.5 Anthropic | 80%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-64kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.0 | 100% confidence 100 percent, Full |
| 104 | Inkling thinkingmachines | 79.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.5 | 100% confidence 100 percent, Full |
| 105 | Grok 4.5 xAI | 79.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.2 | 100% confidence 100 percent, Full |
| 106 | GPT-5.2 OpenAI | 78.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 78.7 | 100% confidence 100 percent, Full |
| 107 | Inkling Small thinkingmachines | 78%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 78.0 | 100% confidence 100 percent, Full |
| 108 | GPT-5.6 Terra OpenAI | 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.0 | 100% confidence 100 percent, Full |
| 109 | GLM-5.2 Z.ai | 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.0 | 100% confidence 100 percent, Full |
| 110 | Gemini 3.6 Flash Google | 76.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.5 | 100% confidence 100 percent, Full |
| 111 | GPT-5.6 Luna OpenAI | 76.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.5 | 100% confidence 100 percent, Full |
| 112 | GPT-5.5 OpenAI | 76.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.2 | 100% confidence 100 percent, Full |
| 113 | Claude Opus 4.5 Anthropic | 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 114 | Grok 4.6 xAI | 74.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 74.8 | 100% confidence 100 percent, Full |
| 115 | GPT-5.6 Sol OpenAI | 74.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 74.5 | 100% confidence 100 percent, Full |
| 116 | GPT-6 Luna OpenAI | 73%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 73.0 | 100% confidence 100 percent, Full |
| 117 | GPT-5.1 OpenAI | 72.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.8 | 85% confidence 85 percent, High |
| 118 | GPT-5.2 OpenAI | 72.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.7 | 100% confidence 100 percent, Full |
| 119 | GPT-6 Sol OpenAI | 72.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.2 | 100% confidence 100 percent, Full |
| 120 | Claude Opus 4.5 Anthropic | 72%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.0 | 100% confidence 100 percent, Full |
| 121 | GLM-5.3-Flash Z.ai | 71.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 71.8 | 100% confidence 100 percent, Full |
| 122 | GPT-6 Luna OpenAI | 70.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.3 | 100% confidence 100 percent, Full |
| 123 | GPT-5 Pro OpenAI | 70.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.2 | 64% confidence 64 percent, Medium |
| 124 | Qwen3.8 27B Alibaba / Qwen | 69.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 69.2 | 69% confidence 69 percent, Medium |
| 125 | Qwen3.8 27B Alibaba / Qwen | 68.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.7 | 69% confidence 69 percent, Medium |
| 126 | GPT-5.4 OpenAI | 68.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.2 | 100% confidence 100 percent, Full |
| 127 | Inkling Small thinkingmachines | 67%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.0 | 100% confidence 100 percent, Full |
| 128 | GPT-5 OpenAI | 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.7 | 100% confidence 100 percent, Full |
| 129 | Kimi K3 Moonshot AI | 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.7 | 100% confidence 100 percent, Full |
| 130 | Kimi K2.5 Moonshot AI | 65.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.3 | 100% confidence 100 percent, Full |
| 131 | Claude Sonnet 4.5 (latest) Anthropic | 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.7 | 43% confidence 43 percent, Low |
| 132 | MiniMax-M2.5 minimax | 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.7 | 100% confidence 100 percent, Full |
| 133 | GPT-5.4 mini OpenAI | 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.7 | 100% confidence 100 percent, Full |
| 134 | Grok 4.7 xAI | 63.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.3 | 100% confidence 100 percent, Full |
| 135 | GPT-6 Luna OpenAI | 61%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.0 | 100% confidence 100 percent, Full |
| 136 | o3 OpenAI | 60.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.8 | 87% confidence 87 percent, High |
| 137 | GPT-5.6 Terra OpenAI | 60.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.2 | 100% confidence 100 percent, Full |
| 138 | o3-pro OpenAI | 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.3 | 22% confidence 22 percent, Low |
| 139 | Claude Opus 4.5 Anthropic | 58.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.7 | 100% confidence 100 percent, Full |
| 140 | o4-mini OpenAI | 58.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.7 | 87% confidence 87 percent, High |
| 141 | GPT-5.4 mini OpenAI | 58.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.0 | 100% confidence 100 percent, Full |
| 142 | Gemini 3 Flash Preview Google | 57.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.7 | 93% confidence 93 percent, High |
| 143 | GPT-5.1 OpenAI | 57.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.7 | 85% confidence 85 percent, High |
| 144 | DeepSeek V3.2 DeepSeek | 57.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.0 | 56% confidence 56 percent, Medium |
| 145 | o3-pro OpenAI | 57.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.0 | 22% confidence 22 percent, Low |
| 146 | GPT-5.6 Luna OpenAI | 56.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.5 | 100% confidence 100 percent, Full |
| 147 | GPT-5 OpenAI | 56.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.2 | 100% confidence 100 percent, Full |
| 148 | GPT-5.2 OpenAI | 55.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 55.7 | 100% confidence 100 percent, Full |
| 149 | GPT-5 Mini OpenAI | 54.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.3 | 99% confidence 99 percent, High |
| 150 | o3 OpenAI | 53.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 53.8 | 87% confidence 87 percent, High |
| 151 | Gemini 3.5 Flash Lite Google | 53.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 53.5 | 100% confidence 100 percent, Full |
| 152 | Inkling Small thinkingmachines | 52.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.5 | 100% confidence 100 percent, Full |
| 153 | GPT-5.4 nano OpenAI | 51.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.5 | 100% confidence 100 percent, Full |
| 154 | Claude Sonnet 4.5 (latest) Anthropic | 48.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 48.3 | 43% confidence 43 percent, Low |
| 155 | Claude Haiku 4.5 (latest) Anthropic | 47.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.7 | 33% confidence 33 percent, Low |
| 156 | GLM-5.3-Flash Z.ai | 47.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.0 | 100% confidence 100 percent, Full |
| 157 | Claude Sonnet 4.5 (latest) Anthropic | 46.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.5 | 43% confidence 43 percent, Low |
| 158 | GLM-5 Z.ai | 44.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.7 | 90% confidence 90 percent, High |
| 159 | o3-pro OpenAI | 44.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.3 | 22% confidence 22 percent, Low |
| 160 | GPT-5 OpenAI | 44%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.0 | 100% confidence 100 percent, Full |
| 161 | o4-mini OpenAI | 41.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 41.8 | 87% confidence 87 percent, High |
| 162 | o3 OpenAI | 41.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 41.5 | 87% confidence 87 percent, High |
| 163 | Gemini 2.5 Pro Google | 41%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 41.0 | 90% confidence 90 percent, High |
| 164 | GPT-5.4 mini OpenAI | 40.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 40.8 | 100% confidence 100 percent, Full |
| 165 | Claude Opus 4.5 Anthropic | 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 40.0 | 100% confidence 100 percent, Full |
| 166 | Claude Sonnet 4 Anthropic | 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 40.0 | 100% confidence 100 percent, Full |
| 167 | GPT-5.4 nano OpenAI | 38.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.2 | 100% confidence 100 percent, Full |
| 168 | GPT-6 Luna OpenAI | 37.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.7 | 100% confidence 100 percent, Full |
| 169 | Claude Haiku 4.5 (latest) Anthropic | 37.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.3 | 33% confidence 33 percent, Low |
| 170 | GPT-5 Mini OpenAI | 37.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.3 | 99% confidence 99 percent, High |
| 171 | Gemini 2.5 Pro Google | 37%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.0 | 90% confidence 90 percent, High |
| 172 | Claude Opus 4 Anthropic | 35.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 35.7 | 100% confidence 100 percent, Full |
| 173 | o3-mini OpenAI | 34.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 34.5 | 96% confidence 96 percent, High |
| 174 | GPT-5.6 Luna OpenAI | 34.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 34.2 | 100% confidence 100 percent, Full |
| 175 | GPT-5.1 OpenAI | 33.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.2 | 85% confidence 85 percent, High |
| 176 | GPT-5.4 nano OpenAI | 33%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.0 | 100% confidence 100 percent, Full |
| 177 | Gemini 3.5 Flash Lite Google | 32.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 32.3 | 100% confidence 100 percent, Full |
| 178 | Claude Sonnet 4.5 (latest) Anthropic | 31%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.0 | 43% confidence 43 percent, Low |
| 179 | Claude Opus 4 Anthropic | 30.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.7 | 100% confidence 100 percent, Full |
| 180 | Gemini 2.5 Pro Google | 29.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.5 | 90% confidence 90 percent, High |
| 181 | Claude Sonnet 4 Anthropic | 29.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.0 | 100% confidence 100 percent, Full |
| 182 | Gemini 3 Flash Preview Google | 29.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.0 | 93% confidence 93 percent, High |
| 183 | Claude Sonnet 3.7 Anthropic | 28.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 28.6 | 100% confidence 100 percent, Full |
| 184 | Claude Sonnet 4 Anthropic | 28.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 28.0 | 100% confidence 100 percent, Full |
| 185 | Claude Opus 4 Anthropic | 27%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.0 | 100% confidence 100 percent, Full |
| 186 | GPT-5 Mini OpenAI | 26.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 26.3 | 99% confidence 99 percent, High |
| 187 | Claude Haiku 4.5 (latest) Anthropic | 25.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.5 | 33% confidence 33 percent, Low |
| 188 | Claude Sonnet 4.5 (latest) Anthropic | 25.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.5 | 43% confidence 43 percent, Low |
| 189 | Claude Sonnet 4 Anthropic | 23.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 23.8 | 100% confidence 100 percent, Full |
| 190 | Claude Opus 4 Anthropic | 22.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 22.5 | 100% confidence 100 percent, Full |
| 191 | o3-mini OpenAI | 22.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 22.3 | 96% confidence 96 percent, High |
| 192 | o4-mini OpenAI | 21.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 21.3 | 87% confidence 87 percent, High |
| 193 | DeepSeek-R1 DeepSeek | 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 21.2 | 88% confidence 88 percent, High |
| 194 | Claude Sonnet 3.7 Anthropic | 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 21.2 | 100% confidence 100 percent, Full |
| 195 | GPT-5 Nano OpenAI | 20.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 20.7 | 85% confidence 85 percent, High |
| 196 | GPT-5.4 nano OpenAI | 18.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.3 | 100% confidence 100 percent, Full |
| 197 | Gemini 3.5 Flash Lite Google | 17%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.0 | 100% confidence 100 percent, Full |
| 198 | Claude Haiku 4.5 (latest) Anthropic | 16.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.8 | 33% confidence 33 percent, Low |
| 199 | GPT-5 Nano OpenAI | 16.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.7 | 85% confidence 85 percent, High |
| 200 | Gemini 2.5 Pro Google | 16%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.0 | 90% confidence 90 percent, High |
| 201 | DeepSeek-R1 DeepSeek | 15.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] R1Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 15.8 | 88% confidence 88 percent, High |
| 202 | o3-mini OpenAI | 14.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 14.5 | 96% confidence 96 percent, High |
| 203 | Claude Haiku 4.5 (latest) Anthropic | 14.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 14.3 | 33% confidence 33 percent, Low |
| 204 | Claude Sonnet 3.7 Anthropic | 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.6 | 100% confidence 100 percent, Full |
| 205 | GPT-5.4 mini OpenAI | 13%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.0 | 100% confidence 100 percent, Full |
| 206 | GPT-5.2 OpenAI | 12.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 12.3 | 100% confidence 100 percent, Full |
| 207 | Claude Sonnet 3.7 Anthropic | 11.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 11.6 | 100% confidence 100 percent, Full |
| 208 | Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen | 11%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 11.0 | 48% confidence 48 percent, Low |
| 209 | Magistral Medium (latest) mistral | 6.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.1 | 16% confidence 16 percent, Low |
| 210 | Magistral Medium (latest) mistral | 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.9 | 16% confidence 16 percent, Low |
| 211 | GPT-5.1 OpenAI | 5.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.8 | 85% confidence 85 percent, High |
| 212 | GPT-4.1 OpenAI | 5.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.5 | 100% confidence 100 percent, Full |
| 213 | Magistral Small mistral | 5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.0 | 48% confidence 48 percent, Low |
| 214 | GPT-4o OpenAI | 4.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-2024-11-20Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.5 | 59% confidence 59 percent, Medium |
| 215 | Llama 4 Maverick 17B Instruct Meta | 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.4 | 78% confidence 78 percent, Medium |
| 216 | GPT-5 Nano OpenAI | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 85% confidence 85 percent, High |
| 217 | GPT-4.1 mini OpenAI | 3.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.5 | 87% confidence 87 percent, High |
| 218 | GPT-4.1 nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 82% confidence 82 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.