Conceptualized
The bench is specified: synthetic vendors that run live in the task, a ticket plus starter repo plus vendor docs as the only inputs, and a grader that reads the request log and store state rather than the agent's code. Two sets are declared: a public development set and an unpublished held-out set.
What shipped
- Scope. A benchmark for coding agents on third-party API integrations: read the vendor docs, ship the connector, score the run from the traffic.
- Vendors. Synthetic, generated from a declaration, live inside the task. Real vendor APIs are out: they change, cannot be reset, and already sit in training data.
- Inputs. Ticket (
PROBLEM.md), starter repository, vendor documentation. The agent does not get the answer key, the reference connector, or the vendor declaration. - Grading. After the agent stops: data-store correctness, protocol from the vendor request log, fault survival. Source is not scored.
- Sets. A public development set and an unpublished held-out set. Held-out results are the capability number.