Signal fires your API, chains the values it returns, then verifies what actually landed in Postgres, MySQL, Redis and Mongo.
The same harness measures what answers differently every time — an LLM step, an agent's tool choice, a reply graded against a rubric — as a pass rate instead of a single green run.
That is where the interesting bugs start. A webhook that never fires, a ledger row that never lands, a cache that never invalidates — all of it returns 200.
Collections, folders, saved examples, variable highlighting and code generation. Import a Postman collection or an OpenAPI 3 spec and keep working.
Extract a value, assert on it, feed it into the next step. Conditions, loops, retries and sub-flows called by reference.
rowCount, row values, Redis keys, Mongo documents — asserted in the same run as the request that caused them.
Describe the scenario and Claude drafts the flow. When a run goes red it names the root cause and the fix — on the run, suite and trend pages, and even inside the Slack message.
Stub the APIs you depend on: method + path in, your canned answer out. Fresh {{$guid}} per hit, simulated latency, zero external flakiness.
Group flows into suites, hand them to cron, and fan independent flows out across workers. Flaky tests get quarantined automatically instead of burying the suite.
Results land in Slack or any webhook — standing rules per workspace, test or suite. Failure messages can carry Claude's analysis along.
One POST runs a suite and answers with JUnit XML. The README badge and a public status page are fed by real test runs.
Companies, workspaces, members and API tokens — every resource scoped, including the MCP session. Google/GitHub sign-in included.
Point the same harness at systems that answer differently every time: an LLM step, a judge operator that grades a reply against a plain-language rubric and records its reason, and an agent step that captures which tools the model actually called — toolSequence == search › book is an ordinary assertion. Repeat a row N times and read a pass rate instead of guessing.
One authenticated call runs a flow and answers with the verdict — 200 when it passes, 422 when it does not. Or let cron, your pipeline or Claude start it.