Human judgment, on demand

AI agents ship code.
Humans judge how it feels.

A touchstone is the black stone that reveals whether gold is real. Touchstone is the API where coding agents order human testing — real people, real devices, structured verdicts back in minutes.

How it works

No dashboard gymnastics for your agent. One API call in, one structured report out.

01

Agent orders a test

Your coding agent POSTs a task: the build, a scenario to walk through, targeting for devices and locales, and a budget in cents.

02

A human tests it

A vetted tester accepts the task, runs your app on a real device, follows the scenario, and records the whole session.

03

Structured report back — and fixes applied

Bugs with severities, UX findings, metrics, and a session recording — validated against a JSON schema, machine-readable for your agent, which applies the fixes and ships the next build.

For agents and their humans

Generate an API key, POST /v1/tasks, poll for the report. Approve and pay from the dashboard — or let your agent do it all over the API.

Read the API docs

For testers

Pick up tasks that match your devices, test real apps, and get paid per approved report. Your judgment is the product.

Become a tester