Skip to content

Models on Integration Bench

Every model evaluated on Integration Bench rev 01 with its per-task results.

Claude Fable 5 on Integration Bench Claude Fable 5 scored 67.68 on Integration Bench rev 01, resolving 34 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Claude Haiku 4.5 on Integration Bench Claude Haiku 4.5 scored 15.10 on Integration Bench rev 01, resolving 8 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Claude Opus 4.8 on Integration Bench Claude Opus 4.8 scored 55.58 on Integration Bench rev 01, resolving 28 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Claude Opus 5 on Integration Bench Claude Opus 5 scored 59.32 on Integration Bench rev 01, resolving 30 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Claude Sonnet 5 on Integration Bench Claude Sonnet 5 scored 53.85 on Integration Bench rev 01, resolving 27 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. DeepSeek V4 Flash (max) on Integration Bench DeepSeek V4 Flash (max) scored 55.68 on Integration Bench rev 01, resolving 28 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. DeepSeek V4 Pro (max) on Integration Bench DeepSeek V4 Pro (max) scored 63.35 on Integration Bench rev 01, resolving 32 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Gemini 3.7 Flash on Integration Bench Gemini 3.7 Flash scored 35.85 on Integration Bench rev 01, resolving 18 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. GLM 5.2 (xhigh) on Integration Bench GLM 5.2 (xhigh) scored 49.68 on Integration Bench rev 01, resolving 25 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. GPT-5.6 Luna on Integration Bench GPT-5.6 Luna scored 51.78 on Integration Bench rev 01, resolving 26 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. GPT-5.6 Sol on Integration Bench GPT-5.6 Sol scored 63.62 on Integration Bench rev 01, resolving 32 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. GPT-5.6 Terra on Integration Bench GPT-5.6 Terra scored 53.79 on Integration Bench rev 01, resolving 27 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Grok 4.6 on Integration Bench Grok 4.6 scored 61.18 on Integration Bench rev 01, resolving 31 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Kimi K3 on Integration Bench Kimi K3 scored 57.47 on Integration Bench rev 01, resolving 29 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Muse Spark 1.2 on Integration Bench Muse Spark 1.2 scored 53.70 on Integration Bench rev 01, resolving 27 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Qwen 3.8 2.4T A95B on Integration Bench Qwen 3.8 2.4T A95B scored 61.68 on Integration Bench rev 01, resolving 31 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory. Qwen 3.8 27B (xhigh) on Integration Bench Qwen 3.8 27B (xhigh) scored 57.73 on Integration Bench rev 01, resolving 29 of 50 public integration tasks. Per-task scores, cost, tool calls and a link to every trajectory.