Benchmarks

The catalog, scored by category.

Sample bars for layout preview. These figures are not measured results.

Reasoning

Multi-step logic and hard questions.

Coding

Patches, terminal tasks, and repo work.

Agentic

Tool use and multi-step jobs.

Knowledge

Facts, instruction following, and recall.

Overall

Same sample scores, shown together. A stronger green is a higher score.

AI modelCompositeReasoningCodingAgenticKnowledge
01
86.4
88.2
85.1
87.4
84.9
02
82.7
83.4
84.8
80.2
82.4
03
79.8
81.0
78.4
80.1
79.7
04
76.5
77.5
74.2
79.6
74.7
05
73.9
73.8
76.1
72.4
73.3
06
71.2
72.0
69.4
70.8
72.6
07
67.4
62.1
58.4
64.0
61.2
08
63.8
55.2
51.8
60.4
54.1