Claude Fable 5.1 and Gemini 3.8 Flash added
Two models join the rev 01 leaderboard. Claude Fable 5.1 scores 69.35 and takes first place; Gemini 3.8 Flash scores 59.68 and lands seventh. All 100 attempts are published in full, same scorer, same 50 public tasks.
What shipped
- Claude Fable 5.1 via Claude Code at
xhigh: 69.35, 35 of 50 tasks resolved, $322.99 for the sweep. First place, 1.67 points ahead of Claude Fable 5. By surface: polling 75.61, writeback 70.56, webhooks 43.92. Model page. - Gemini 3.8 Flash via OpenCode at
high: 59.68, 30 of 50 resolved, seventh place. Polling 63.91, writeback 61.55, webhooks 27.27. Metered at $0 through the provider; cost columns show a dash. Model page. - Data. 100 new attempt bundles at data.integrationbench.com, same layout as the rest:
summary.json,trajectory.json,verdict.json,vendor-requests.json,patch.diff, rawtranscript.jsonl,problem.md. The leaderboard, matrix and per-task indexes are regenerated; the 850 existing bundles are byte-identical.
Notes
- Same task set, same scorer (
task-score-v4-mandatory-gated), one official attempt per task, 60-minute wall clock. Existing rows do not change; ranks below the new entries shift down by one or two. - Seven Fable 5.1 attempts (tasks 0001, 0002, 0014, 0017, 0022, 0026 and 0045) were graded a second time after verifier fixes. The transcript, diff and vendor log are the original run’s; the verdict is from the regrade. Cost, duration and token counts for those seven are read from the harness transcript rather than the grader’s record.
- The webhook gap holds. Fable 5.1 drops 32 points from polling to webhook tasks and Gemini 3.8 Flash drops 37, in line with every other row on the board.