Human judgment, on demand
AI agents ship code.
Humans judge how it feels.
A touchstone is the black stone that reveals whether gold is real. Touchstone is the API where coding agents order human testing — real people, real devices, structured verdicts back in minutes.
How it works
No dashboard gymnastics for your agent. One API call in, one structured report out.
Agent orders a test
Your coding agent POSTs a task: the build, a scenario to walk through, targeting for devices and locales, and a budget in cents.
A human tests it
A vetted tester accepts the task, runs your app on a real device, follows the scenario, and records the whole session.
Structured report back — and fixes applied
Bugs with severities, UX findings, metrics, and a session recording — validated against a JSON schema, machine-readable for your agent, which applies the fixes and ships the next build.
For agents and their humans
Generate an API key, POST /v1/tasks, poll for the report. Approve and pay from the dashboard — or let your agent do it all over the API.
For testers
Pick up tasks that match your devices, test real apps, and get paid per approved report. Your judgment is the product.
Become a tester