livebench coding code completion 2026_06_25 Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 2 | GPT-5.2 Codex OpenAI | 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 87.0 | 43% confidence 43 percent, Low |
| 3 | GPT-5.6 Luna OpenAI | 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 4 | Claude Opus 4.8 Anthropic | 84.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 84.8 | 100% confidence 100 percent, Full |
| 5 | Claude Sonnet 5.5 Anthropic | 84.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 84.8 | 93% confidence 93 percent, High |
| 6 | GPT-5.6 Sol OpenAI | 84.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 84.8 | 100% confidence 100 percent, Full |
| 7 | Claude Fable 5.1 Anthropic | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 8 | Claude Opus 5 Anthropic | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 9 | DeepSeek V4.1 Flash DeepSeek | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 69% confidence 69 percent, Medium |
| 10 | Kimi K3 Moonshot AI | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 11 | GPT-5.5 OpenAI | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 12 | GPT-6.1 Sol OpenAI | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 13 | Claude Fable 5 Anthropic | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 14 | Claude Opus 4.5 Anthropic | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 15 | Muse Spark 1.2 Meta | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 53% confidence 53 percent, Medium |
| 16 | Muse Spark 1.3 Meta | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 93% confidence 93 percent, High |
| 17 | GPT-5.4 OpenAI | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 18 | GPT-5.6 Terra OpenAI | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 19 | GPT-6 Astra OpenAI | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 20 | GPT-6 Luna OpenAI | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 21 | GPT-6 Sol OpenAI | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 22 | GLM-5.2 Z.ai | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 23 | GLM-5.3 Z.ai | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 93% confidence 93 percent, High |
| 24 | GLM-5.3-Flash Z.ai | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 25 | Claude Haiku 5.5 Anthropic | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 26 | Claude Opus 4.7 Anthropic | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 27 | Claude Sonnet 4.6 Anthropic | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 28 | Claude Sonnet 5 Anthropic | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 93% confidence 93 percent, High |
| 29 | DeepSeek V4 Pro 0813 DeepSeek | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 69% confidence 69 percent, Medium |
| 30 | Gemini 3.1 Pro Preview Google | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 31 | Gemini 3.6 Flash Google | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 32 | Muse Spark 1.1 Meta | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 61% confidence 61 percent, Medium |
| 33 | Mistral Large 4 mistral | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 32% confidence 32 percent, Low |
| 34 | Kimi K2.6 Moonshot AI | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 85% confidence 85 percent, High |
| 35 | Grok 4.7 xAI | 78.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 36 | Qwen3.6 Plus Alibaba / Qwen | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 80% confidence 80 percent, High |
| 37 | Qwen3.8 Flash Next Alibaba / Qwen | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 21% confidence 21 percent, Low |
| 38 | Claude Opus 4.6 Anthropic | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 39 | Gemini 3.5 Flash Google | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 40 | Gemini 3.5 Flash Lite Google | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 41 | Gemini 3.7 Flash Google | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 42 | Kimi K2.7 Code Moonshot AI | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 48% confidence 48 percent, Low |
| 43 | GPT-5.2 OpenAI | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 44 | Grok 4.6 xAI | 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 76.1 | 100% confidence 100 percent, Full |
| 45 | GPT-5.4 mini OpenAI | 75.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 75.6 | 100% confidence 100 percent, Full |
| 46 | Qwen3.8 27B Alibaba / Qwen | 73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 73.9 | 69% confidence 69 percent, Medium |
| 47 | Qwen3.8 Max Alibaba / Qwen | 73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 73.9 | 85% confidence 85 percent, High |
| 48 | DeepSeek V4 Flash 0731 DeepSeek | 73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 73.9 | 69% confidence 69 percent, Medium |
| 49 | Qwen3.6 27B Alibaba / Qwen | 71.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 71.7 | 53% confidence 53 percent, Medium |
| 50 | Gemini 3.8 Flash Google | 71.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 71.7 | 100% confidence 100 percent, Full |
| 51 | Qwen3.7 Max Alibaba / Qwen | 69.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 69.6 | 53% confidence 53 percent, Medium |
| 52 | DeepSeek V4 Pro DeepSeek | 69.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 69.6 | 85% confidence 85 percent, High |
| 53 | Nemotron 3 Ultra 550B A55B nvidia | 69.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 69.6 | 100% confidence 100 percent, Full |
| 54 | Grok 4.5 xAI | 69.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 69.6 | 100% confidence 100 percent, Full |
| 55 | DeepSeek V4 Flash Vision Exp DeepSeek | 67.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 67.4 | 21% confidence 21 percent, Low |
| 56 | MiniMax-M3 minimax | 67.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 67.4 | 100% confidence 100 percent, Full |
| 57 | GPT-5.4 nano OpenAI | 67.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 67.4 | 100% confidence 100 percent, Full |
| 58 | Inkling thinkingmachines | 67.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 67.4 | 100% confidence 100 percent, Full |
| 59 | Grok Build 0.1 xAI | 67.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 67.4 | 16% confidence 16 percent, Low |
| 60 | DeepSeek V4 Flash DeepSeek | 65.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 65.2 | 53% confidence 53 percent, Medium |
| 61 | Grok 4.3 xAI | 65.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ | 65.2 | 80% confidence 80 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.