Skip to content
  1. Home
  2. Benchmarks
  3. Integration Bench
  4. GPT-5.6 Luna

GPT-5.6 Luna

Rank #14 of 17 Task score 51.78 Resolved 52% of tasks Harness Codex CLI Effort xhigh Mean time 10 min Mean tool calls 38
54
polling
55% of tasks touching this surface resolved
57
writeback
57% of tasks touching this surface resolved
9
webhooks
9% of tasks touching this surface resolved

Published attempts

10 of 50 scored tasks released

One official run per task. Click a row for the full trajectory: every tool call, vendor request, graded check and the diff.