math level 5 Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-5 OpenAI | 98.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.1 | 100% confidence 100 percent, Full |
| 2 | GPT-5 OpenAI | 97.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 3 | GPT-5 Mini OpenAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 99% confidence 99 percent, High |
| 4 | o4-mini OpenAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 87% confidence 87 percent, High |
| 5 | o3 OpenAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 87% confidence 87 percent, High |
| 6 | Claude Sonnet 4.5 Anthropic | 97.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 21, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.7 | 80% confidence 80 percent, High |
| 7 | Qwen3 Max Alibaba / Qwen | 97.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 9, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.1 | 69% confidence 69 percent, Medium |
| 8 | GPT-5 Mini OpenAI | 96.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.8 | 99% confidence 99 percent, High |
| 9 | DeepSeek-R1 DeepSeek | 96.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.6 | 88% confidence 88 percent, High |
| 10 | Claude Haiku 4.5 Anthropic | 96.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.4 | 64% confidence 64 percent, Medium |
| 11 | GPT-5 Nano OpenAI | 95.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.2 | 85% confidence 85 percent, High |
| 12 | GPT-5 Nano OpenAI | 94.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 20, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.9 | 85% confidence 85 percent, High |
| 13 | Claude Sonnet 3.7 Anthropic | 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.2 | 100% confidence 100 percent, Full |
| 14 | Claude Sonnet 3.7 Anthropic | 90.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 12, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 15 | GPT-4.1 mini OpenAI | 87.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.3 | 87% confidence 87 percent, High |
| 16 | Claude Haiku 4.5 Anthropic | 86.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.9 | 64% confidence 64 percent, Medium |
| 17 | Claude Sonnet 3.7 Anthropic | 86.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.3 | 100% confidence 100 percent, Full |
| 18 | Claude Opus 4 Anthropic | 85.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.0 | 100% confidence 100 percent, Full |
| 19 | Claude Sonnet 4 Anthropic | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 100% confidence 100 percent, Full |
| 20 | GPT-4.1 OpenAI | 83.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.0 | 100% confidence 100 percent, Full |
| 21 | Mistral Medium 3 mistral | 81.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.6 | 85% confidence 85 percent, High |
| 22 | DeepSeek V3 0324 DeepSeek | 75.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 1, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.5 | 74% confidence 74 percent, Medium |
| 23 | Gemma 3 27B IT Google | 74.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.0 | 74% confidence 74 percent, Medium |
| 24 | Llama 4 Maverick 17B Instruct Meta | 73.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.0 | 78% confidence 78 percent, Medium |
| 25 | GPT-4.1 nano OpenAI | 70.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 70.0 | 82% confidence 82 percent, High |
| 26 | Qwen3 235B-A22B Alibaba / Qwen | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 3, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 76% confidence 76 percent, Medium |
| 27 | Claude Sonnet 3.7 Anthropic | 68.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.2 | 100% confidence 100 percent, Full |
| 28 | DeepSeek-V3 DeepSeek | 64.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.9 | 93% confidence 93 percent, High |
| 29 | Llama 4 Scout 17B Instruct Meta | 62.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 62.3 | 64% confidence 64 percent, Medium |
| 30 | Claude Sonnet 3.5 v2 Anthropic | 56.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.9 | 100% confidence 100 percent, Full |
| 31 | Qwen Turbo Alibaba / Qwen | 56.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.2 | 47% confidence 47 percent, Low |
| 32 | GPT-4o (2024-08-06) OpenAI | 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.3 | 100% confidence 100 percent, Full |
| 33 | GPT-4o mini OpenAI | 52.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 52.6 | 100% confidence 100 percent, Full |
| 34 | GPT-4o (2024-05-13) OpenAI | 51.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.0 | 93% confidence 93 percent, High |
| 35 | Mistral Large 2.1 mistral | 50.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 50.3 | 100% confidence 100 percent, Full |
| 36 | GPT-4o (2024-11-20) OpenAI | 49.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 49.8 | 61% confidence 61 percent, Medium |
| 37 | Claude Haiku 3.5 Anthropic | 46.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 12, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.4 | 100% confidence 100 percent, Full |
| 38 | Llama-3.3-70B-Instruct Meta | 41.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 41.6 | 100% confidence 100 percent, Full |
| 39 | Llama-3.1-70B-Instruct Meta | 36.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 36.7 | 93% confidence 93 percent, High |
| 40 | Llama-3.1-8B-Instruct Meta | 22.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.9 | 93% confidence 93 percent, High |
| 41 | Claude Haiku 3 Anthropic | 14.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 14.9 | 93% confidence 93 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.