Tasks authored
The 50 public tasks are authored against live synthetic vendors: applicant tracking, HRIS, and job boards; build, fix, harden, and migrate. Each task is internally evaluated, then verifiers and scoring are iterated: gold and floor probes, check coverage, request-log assertions.
What shipped
- 50 public tasks. Against 14 synthetic ATS, HRIS, and job-board vendors. Surfaces: polling on most tickets, writeback on fewer, webhooks on about one in five. A ticket can declare more than one.
- Kinds. Build, fix, harden, migrate. The agent gets a ticket, a starter repo, and vendor docs, including the parts that are wrong on purpose.
- Internal evaluation. Each task is run against known connectors, then revised.
- Verifiers. Assertions over the canonical store and the vendor request log: credentials, retries, pagination, webhook signatures, field values. Rewritten where coverage was thin.
- Scoring. Required checks gate the task; credit for the rest is awarded after that. Iterated on the same gold and floor probes until the suite is internally consistent.