integration bench rev 01 · mean task score · public task set
#ModelTask scoreScore barResolvedStatus
1 Claude Fable 567.6868%
2 GPT-5.6 Sol63.6264%
3 DeepSeek V4 Pro (max)63.3564%
4 Qwen 3.8 2.4T A95B61.6862%
5 Grok 4.661.1862%
6 Claude Opus 559.3260%
7 Qwen 3.8 27B (xhigh)57.7358%
8 Kimi K357.4758%
9 DeepSeek V4 Flash (max)55.6856%
10 Claude Opus 4.855.5856%
11 Claude Sonnet 553.8554%
12 GPT-5.6 Terra53.7954%
13 Muse Spark 1.253.7054%
14 GPT-5.6 Luna51.7852%
15 GLM 5.2 (xhigh)49.6850%
16 Claude Haiku 4.515.1016%
17 Gemini 3.7 Flash35.8536%28/50 graded
This is the short version. Every number above comes from a recorded run. 170 of them are published, each with every tool call, every vendor request and the diff. Replay them at integrationbench.com →
Scorer task-score-v4-mandatory-gated · data 2 Sep 2026 A research project from HeyMilo