hle Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Fable 5.1 Anthropic 60.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; production safeguards with fallbackPublished Sep 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
60.9 100% confidence 100 percent, Full
2 Claude Fable 5 Anthropic 59%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
59.0 100% confidence 100 percent, Full
3 Claude Opus 5 Anthropic 56.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
56.3 100% confidence 100 percent, Full
4 Fugu Ultra sakana 50%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
50.0 100% confidence 100 percent, Full
5 Claude Opus 4.8 Anthropic 49.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
49.8 100% confidence 100 percent, Full
6 Fugu sakana 47.2%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
47.2 100% confidence 100 percent, Full
7 Claude Opus 4.7 Anthropic 46.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
46.9 100% confidence 100 percent, Full
8 Claude Haiku 5.5 Anthropic 45.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
45.9 100% confidence 100 percent, Full
9 Qwen3.8 Max Alibaba / Qwen 43.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
43.6 85% confidence 85 percent, High
10 Qwen3.8 Max Preview Alibaba / Qwen 43.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh, no toolsPublished Aug 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
43.6 6% confidence 6 percent, Low
11 GPT-5.5 Pro OpenAI 43.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
43.1 53% confidence 53 percent, Medium
12 DeepSeek V4 Pro 0813 DeepSeek 42.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
42.7 69% confidence 69 percent, Medium
13 GPT-5.4 Pro OpenAI 42.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
42.7 69% confidence 69 percent, Medium
14 Qwen3.7 Max Alibaba / Qwen 41.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
41.4 53% confidence 53 percent, Medium
15 GPT-5.5 OpenAI 41.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
41.4 100% confidence 100 percent, Full
16 GPT-5.4 OpenAI 39.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
39.8 100% confidence 100 percent, Full
17 DeepSeek V4 Pro DeepSeek 37.7%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
37.7 85% confidence 85 percent, High
18 Qwen3.8 Flash Next Alibaba / Qwen 35.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; GPT-4o judge Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
35.9 21% confidence 21 percent, Low
19 DeepSeek V4 Flash DeepSeek 34.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
34.8 53% confidence 53 percent, Medium
20 Claude Sonnet 4.6 Anthropic 34.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
34.6 100% confidence 100 percent, Full
21 Qwen3.8 27B Alibaba / Qwen 30.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; GPT-4o judge Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
30.8 69% confidence 69 percent, Medium
22 GPT-5.4 mini OpenAI 28.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] without toolsPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
28.2 100% confidence 100 percent, Full
23 Nemotron 3 Ultra 550B A55B nvidia 26.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 4, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
26.7 100% confidence 100 percent, Full
24 GPT-5.4 nano OpenAI 24.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] without toolsPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
24.3 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.