Score intervals on every row; two models retired
Each leaderboard row now carries a lower and upper bound around its Pass 1 score, replacing the placeholder ± figure. Claude Haiku 4.5 and Gemini 3.7 Flash leave the board; the leaderboard is 17 models.
What changed
- Intervals. Every row on the leaderboard shows a ± beside its Pass 1 score, and the error bar on the chart is drawn from the lower bound to the upper. The ± is half the interval width; hover it for the exact bounds, which are per model and not symmetric around the score: Claude Fable 5.1 sits at 69.35 within 67.47 to 71.23, Claude Opus 4.8 at 55.58 within 54.46 to 58.02. This replaces the ± band that had been a placeholder since rev 01.
- Naming. The score column is labelled Pass 1: the Task Score of the one published attempt per task, averaged over the task set. It is the same number the model, task and matrix pages have always shown.
- Data.
leaderboard.jsonrows gainscore_loandscore_hi(absolute scores, not offsets).score_seis gone. Anything reading the old field should switch. - Two rows retired. Claude Haiku 4.5 (15.10) and Gemini 3.7 Flash (35.85, 22 provider failures) are removed from the leaderboard, matrix and per-task pages, and their attempt bundles are removed from data.integrationbench.com. Neither is a model anyone would pick for this work, and both sat far enough below the field to compress the rest of the chart. The rev 01 launch note still records their original scores.
What did not change
- Scores, resolved counts, costs and ranks for the remaining 17 models are identical to the previous update. Same tasks, same scorer, same one official attempt per cell.