Skip to content
  1. Home
  2. Releases
rev 01

rev 01: 17 models on the public tasks

First public leaderboard. Seventeen models, one official attempt on each public task, each run in its vendor's own harness at the highest reasoning effort exposed. Scored with task-score-v4-mandatory-gated.

20 August 2026leaderboard

What shipped

  • Models (17). Claude Opus 5, Opus 4.8, Sonnet 5, Fable 5 and Haiku 4.5 via Claude Code; GPT-5.6 Sol, Luna and Terra via Codex CLI; DeepSeek V4 Pro and Flash, Gemini 3.7 Flash, Grok 4.6, GLM 5.2, Kimi K3, Qwen 3.8 27B, Qwen 3.8 2.4T A95B and Muse Spark 1.2 via OpenCode over OpenRouter. Reasoning effort is set to the highest level the provider exposes (xhigh where available) and recorded per attempt.
  • Tasks. The public task set against 15 invented applicant-tracking, HRIS and job-board vendors: build tickets make up just over half, with the rest split across fix, harden and migrate. By the surface each ticket declares, most touch polling, fewer than half writeback, and about one in five webhooks (non-exclusive).
  • Runs. One official attempt per model per task, 60-minute wall clock each. Trajectories are published task batch by task batch; see the data release. Superseded retries are not part of the published score. Task tree and harness tree hashes, vendor image digests and scorer version are recorded in each attempt’s summary.json.
  • Scorer. task-score-v4-mandatory-gated. Leaderboard metric is mean task score over the full public set with a missing verdict counted as 0.

Caveats recorded

  • Gemini 3.7 Flash: 22 attempts have no verdict because of provider failures; they are baked into the published score as zeros. The model’s row is flagged on the leaderboard.
  • Codex CLI does not report cost, so GPT-5.6 rows and Kimi K3 show no cost figure.
  • Harness differs by model family, which is a confound on every cross-family comparison.
  • Mandatory gating produces a near-binary score distribution on many tasks; read the per-task pages before drawing conclusions from a two-point difference.

The announcement was the RL Supply post Introducing Integration Bench.