hle scale Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 54.8%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 9, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.8 | 100% confidence 100 percent, Full |
| 2 | Claude Fable 5.1 Anthropic | 46.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.5 | 100% confidence 100 percent, Full |
| 3 | Gemini 3.1 Pro Preview Google | 46.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.4 | 100% confidence 100 percent, Full |
| 4 | Gemini 3.8 Flash Google | 44.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 9, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.5 | 100% confidence 100 percent, Full |
| 5 | GPT-5.4 Pro OpenAI | 44.3%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 23, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.3 | 69% confidence 69 percent, Medium |
| 6 | Gemini 3 Pro Preview Google | 37.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 19, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.5 | 68% confidence 68 percent, Medium |
| 7 | GPT-5.4 OpenAI | 36.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 10, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.2 | 100% confidence 100 percent, Full |
| 8 | Claude Opus 4.7 Anthropic | 36.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 22, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.2 | 100% confidence 100 percent, Full |
| 9 | GPT-5 Pro OpenAI | 31.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 6, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.6 | 64% confidence 64 percent, Medium |
| 10 | Claude Haiku 5.5 Anthropic | 30.1%Humanity’s Last ExamPublished steward score [variant] claude-haiku-5-5Published Oct 8, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.1 | 100% confidence 100 percent, Full |
| 11 | GPT-5.2 OpenAI | 27.8%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Dec 15, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.8 | 100% confidence 100 percent, Full |
| 12 | GPT-5 OpenAI | 25.3%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at reasoning_effort: 'high'. [variant] Published Aug 7, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.3 | 100% confidence 100 percent, Full |
| 13 | Kimi K2.5 Moonshot AI | 24.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Feb 13, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 24.4 | 100% confidence 100 percent, Full |
| 14 | GPT-5 Mini OpenAI | 19.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Aug 22, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 19.4 | 99% confidence 99 percent, High |
| 15 | Claude Opus 4.6 Anthropic | 19%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Feb 17, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 19.0 | 100% confidence 100 percent, Full |
| 16 | Claude Opus 4.5 Anthropic | 14.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 26, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 14.2 | 100% confidence 100 percent, Full |
| 17 | Claude Opus 4 Anthropic | 10.7%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 24, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 10.7 | 100% confidence 100 percent, Full |
| 18 | Gemini 3.1 Flash Lite Preview Google | 8.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 23, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.6 | 48% confidence 48 percent, Low |
| 19 | o1-pro OpenAI | 8.1%Humanity’s Last Exam9% (216 prompts) failed due to a post-training bug and were counted as failures. OpenAI has been informed and is working on a fix. --- Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.1 | 16% confidence 16 percent, Low |
| 20 | Claude Sonnet 3.7 Anthropic | 8.0%Humanity’s Last ExamThinking budget: 16,000 tokens. Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.0 | 100% confidence 100 percent, Full |
| 21 | Claude Opus 4.1 Anthropic | 7.9%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Aug 8, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.9 | 80% confidence 80 percent, High |
| 22 | Claude Sonnet 4.5 Anthropic | 7.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Oct 2, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.5 | 80% confidence 80 percent, High |
| 23 | Llama 4 Maverick 17B Instruct Meta | 5.7%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.7 | 78% confidence 78 percent, Medium |
| 24 | Claude Sonnet 4 Anthropic | 5.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 23, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.5 | 100% confidence 100 percent, Full |
| 25 | GPT-4.1 OpenAI | 5.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.4 | 100% confidence 100 percent, Full |
| 26 | Mistral Medium 3 mistral | 4.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 13, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.5 | 85% confidence 85 percent, High |
| 27 | Nova Pro amazon | 4.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.4 | 100% confidence 100 percent, Full |
| 28 | Nova Lite amazon | 3.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.6 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.