Skip to content
  1. Home
  2. Releases
rev 01 · update

Claude Fable 5.1 and Gemini 3.8 Flash added

Two models join the rev 01 leaderboard. Claude Fable 5.1 scores 69.35 and takes first place; Gemini 3.8 Flash scores 59.68 and lands seventh. All 100 attempts are published in full, same scorer, same 50 public tasks.

4 Sep 2026leaderboard

What shipped

  • Claude Fable 5.1 via Claude Code at xhigh: 69.35, 35 of 50 tasks resolved, $322.99 for the sweep. First place, 1.67 points ahead of Claude Fable 5. By surface: polling 75.61, writeback 70.56, webhooks 43.92. Model page.
  • Gemini 3.8 Flash via OpenCode at high: 59.68, 30 of 50 resolved, seventh place. Polling 63.91, writeback 61.55, webhooks 27.27. Metered at $0 through the provider; cost columns show a dash. Model page.
  • Data. 100 new attempt bundles at data.integrationbench.com, same layout as the rest: summary.json, trajectory.json, verdict.json, vendor-requests.json, patch.diff, raw transcript.jsonl, problem.md. The leaderboard, matrix and per-task indexes are regenerated; the 850 existing bundles are byte-identical.

Notes

  • Same task set, same scorer (task-score-v4-mandatory-gated), one official attempt per task, 60-minute wall clock. Existing rows do not change; ranks below the new entries shift down by one or two.
  • Seven Fable 5.1 attempts (tasks 0001, 0002, 0014, 0017, 0022, 0026 and 0045) were graded a second time after verifier fixes. The transcript, diff and vendor log are the original run’s; the verdict is from the regrade. Cost, duration and token counts for those seven are read from the harness transcript rather than the grader’s record.
  • The webhook gap holds. Fable 5.1 drops 32 points from polling to webhook tasks and Gemini 3.8 Flash drops 37, in line with every other row on the board.