AI Benchmarks: Definitions, Weights, and Sources

57 benchmarks feed the SI Score, grouped into 4 capability pillars. Each result on the site links back to the board or dataset it came from.

coding

pillar weight 40%

aider polyglot

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 23 models scored

livebench coding code completion 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench coding code generation 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench coding javascript 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench coding python 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench coding typescript 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

swe bench pro

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 38 models scored

swe bench pro public

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 23 models scored

swe bench pro qwen refined and corrected task set

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

swe bench verified

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 87 models scored

terminal bench

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 28 models scored

terminal bench v0 1

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 3 models scored

terminal bench v2 0

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 9 models scored

terminal bench v2 1

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 41 models scored

terminal bench v3 0

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

terminal bench v4 0

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 28 models scored

reasoning

pillar weight 30%

arc agi 1

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

arc agi 2

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 6 models scored

arc agi 3

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

arc agi v1 public eval

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 64 models scored

arc agi v1 semi private

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 66 models scored

arc agi v2 public eval

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 66 models scored

arc agi v2 semi private

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 68 models scored

arc agi v3 semi private

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 18 models scored

gpqa diamond

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 0.5 within pillar · unit % · 121 models scored

hle

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 24 models scored

hle 1811 verified and revised items

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

hle full set

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

hle full set text mm

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

hle full set tools

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

hle scale

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 22 models scored

hle text only

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

hle text only subset

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 2 models scored

hle text only subset tools

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

hle text only tools

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

hle tools

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 26 models scored

livebench reasoning connections 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench reasoning consecutive events 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench reasoning logic with navigation 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench reasoning spatial 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench reasoning theory of mind 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench reasoning zebra puzzle 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

mmlu pro

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 0.5 within pillar · unit % · 5 models scored

math

pillar weight 15%

frontiermath tier 1 3

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 4 models scored

frontiermath tier 4

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 4 models scored

frontiermath tier 4 v2

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 44 models scored

frontiermath tiers 1 3 v2

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 55 models scored

frontiermath vv2 tier 1 3

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 3 models scored

frontiermath vv2 tier 4

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 1 models scored

livebench math amps hard 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench math integrals with game 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench math math comp 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench math olympiad 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

livebench math simplify 2026_06_25

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 60 models scored

math level 5

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 1 within pillar · unit % · 22 models scored

otis mock aime 2024 2025

Published result; benchmark version and evaluation conditions remain in the id and result note.

weight 0.5 within pillar · unit % · 104 models scored

preference

pillar weight 15%