arc agi v3 semi private Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6 Astra OpenAI 62.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
62.7 100% confidence 100 percent, Full
2 GPT-6 Astra OpenAI 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.3 100% confidence 100 percent, Full
3 GPT-6 Astra OpenAI 54.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
54.8 100% confidence 100 percent, Full
4 GPT-6.1 Sol OpenAI 52.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.7 100% confidence 100 percent, Full
5 GPT-6.1 Sol OpenAI 39.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
39.9 100% confidence 100 percent, Full
6 GPT-6 Astra OpenAI 38.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
38.6 100% confidence 100 percent, Full
7 Claude Opus 5 Anthropic 30.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
30.2 100% confidence 100 percent, Full
8 GPT-6.1 Sol OpenAI 26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
26.7 100% confidence 100 percent, Full
9 GPT-6 Astra OpenAI 17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.5 100% confidence 100 percent, Full
10 GPT-6.1 Sol OpenAI 10.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
10.6 100% confidence 100 percent, Full
11 Gemini 3.8 Flash Google 10.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
10.4 100% confidence 100 percent, Full
12 GPT-5.6 Sol OpenAI 7.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.8 100% confidence 100 percent, Full
13 GPT-5.6 Sol OpenAI 7.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.0 100% confidence 100 percent, Full
14 Gemini 3.8 Flash Google 6.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.0 100% confidence 100 percent, Full
15 GPT-6 Sol OpenAI 4.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.6 100% confidence 100 percent, Full
16 Gemini 3.8 Flash Google 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 100% confidence 100 percent, Full
17 GPT-6.1 Sol OpenAI 3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.9 100% confidence 100 percent, Full
18 GPT-5.6 Sol OpenAI 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.1 100% confidence 100 percent, Full
19 Grok 4.6 xAI 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.1 100% confidence 100 percent, Full
20 Grok 4.7 xAI 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.8 100% confidence 100 percent, Full
21 GPT-6 Sol OpenAI 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.8 100% confidence 100 percent, Full
22 Grok 4.7 xAI 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.7 100% confidence 100 percent, Full
23 Grok 4.7 xAI 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.7 100% confidence 100 percent, Full
24 Claude Opus 4.8 Anthropic 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.5 100% confidence 100 percent, Full
25 GPT-5.6 Sol OpenAI 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.1 100% confidence 100 percent, Full
26 GPT-6 Sol OpenAI 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.9 100% confidence 100 percent, Full
27 GPT-5.6 Terra OpenAI 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.8 100% confidence 100 percent, Full
28 GPT-5.6 Terra OpenAI 0.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.7 100% confidence 100 percent, Full
29 GPT-5.6 Terra OpenAI 0.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.5 100% confidence 100 percent, Full
30 GPT-5.5 OpenAI 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-5-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 100% confidence 100 percent, Full
31 Gemini 3.1 Pro Preview Google 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] google-gemini-3-1-pro-previewPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 100% confidence 100 percent, Full
32 GPT-6 Sol OpenAI 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 100% confidence 100 percent, Full
33 GPT-5.6 Sol OpenAI 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.3 100% confidence 100 percent, Full
34 Grok 4.5 xAI 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.3 100% confidence 100 percent, Full
35 Grok 4.5 xAI 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.3 100% confidence 100 percent, Full
36 Grok 4.5 xAI 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.3 100% confidence 100 percent, Full
37 Grok 4.7 xAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
38 GPT-5.4 OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-4-2026-03-05-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
39 GPT-6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
40 GPT-6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
41 Claude Opus 4.7 Anthropic 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] anthropic-opus-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
42 GPT-5.6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
43 GPT-5.6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
44 GPT-5.6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
45 GPT-6 Luna OpenAI 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.2 100% confidence 100 percent, Full
46 GPT-6 Sol OpenAI 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.1 100% confidence 100 percent, Full
47 GPT-6 Luna OpenAI 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.1 100% confidence 100 percent, Full
48 GPT-5.6 Luna OpenAI 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.1 100% confidence 100 percent, Full
49 Grok 4.20 (Reasoning) xAI 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-20-beta-0309-reasoningPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.1 80% confidence 80 percent, High
50 GPT-5.6 Terra OpenAI 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.1 100% confidence 100 percent, Full
51 GPT-6 Luna OpenAI 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full
52 GPT-5.6 Luna OpenAI 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full
53 GPT-5.6 Terra OpenAI 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.