AI Benchmarks: Definitions, Weights, and Sources
57 benchmarks feed the SI Score, grouped into 4 capability pillars. Each result on the site links back to the board or dataset it came from.
coding
pillar weight 40%aider polyglot
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 23 models scored
livebench coding code completion 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench coding code generation 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench coding javascript 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench coding python 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench coding typescript 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
swe bench pro
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 38 models scored
swe bench pro public
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 23 models scored
swe bench pro qwen refined and corrected task set
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
swe bench verified
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 87 models scored
terminal bench
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 28 models scored
terminal bench v0 1
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 3 models scored
terminal bench v2 0
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 9 models scored
terminal bench v2 1
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 41 models scored
terminal bench v3 0
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
terminal bench v4 0
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 28 models scored
reasoning
pillar weight 30%arc agi 1
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
arc agi 2
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 6 models scored
arc agi 3
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
arc agi v1 public eval
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 64 models scored
arc agi v1 semi private
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 66 models scored
arc agi v2 public eval
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 66 models scored
arc agi v2 semi private
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 68 models scored
arc agi v3 semi private
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 18 models scored
gpqa diamond
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 0.5 within pillar · unit % · 121 models scored
hle
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 24 models scored
hle 1811 verified and revised items
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
hle full set
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
hle full set text mm
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
hle full set tools
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
hle scale
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 22 models scored
hle text only
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
hle text only subset
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 2 models scored
hle text only subset tools
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
hle text only tools
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
hle tools
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 26 models scored
livebench reasoning connections 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench reasoning consecutive events 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench reasoning logic with navigation 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench reasoning spatial 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench reasoning theory of mind 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench reasoning zebra puzzle 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
mmlu pro
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 0.5 within pillar · unit % · 5 models scored
math
pillar weight 15%frontiermath tier 1 3
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 4 models scored
frontiermath tier 4
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 4 models scored
frontiermath tier 4 v2
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 44 models scored
frontiermath tiers 1 3 v2
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 55 models scored
frontiermath vv2 tier 1 3
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 3 models scored
frontiermath vv2 tier 4
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 1 models scored
livebench math amps hard 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench math integrals with game 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench math math comp 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench math olympiad 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
livebench math simplify 2026_06_25
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 60 models scored
math level 5
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 1 within pillar · unit % · 22 models scored
otis mock aime 2024 2025
Published result; benchmark version and evaluation conditions remain in the id and result note.
weight 0.5 within pillar · unit % · 104 models scored