Skip to content
  1. Home
  2. Releases

Conceptualized

The bench is specified: synthetic vendors that run live in the task, a ticket plus starter repo plus vendor docs as the only inputs, and a grader that reads the request log and store state rather than the agent's code. Two sets are declared: a public development set and an unpublished held-out set.

15 July 2026

What shipped

  • Scope. A benchmark for coding agents on third-party API integrations: read the vendor docs, ship the connector, score the run from the traffic.
  • Vendors. Synthetic, generated from a declaration, live inside the task. Real vendor APIs are out: they change, cannot be reset, and already sit in training data.
  • Inputs. Ticket (PROBLEM.md), starter repository, vendor documentation. The agent does not get the answer key, the reference connector, or the vendor declaration.
  • Grading. After the agent stops: data-store correctness, protocol from the vendor request log, fault survival. Source is not scored.
  • Sets. A public development set and an unpublished held-out set. Held-out results are the capability number.