Technical proof
How we prove it works.
An agent that serves customers doesn't get approved by watching a demo. Here's what gets tested before production, what gets watched after, and who signs when something goes off script. Every step is written up in Resources — they're the same documents we use internally.
01
Before production: the scenario bank
A list of hard conversations with their expected outcome, written before we build. It doesn't measure whether the agent answered nicely: it measures pass or fail. The whole bank runs again before any change — instructions, catalog or model — touches production.
Read the full method02
The scenarios that matter are the awkward ones
Out of stock, a size that doesn't exist, a price that changed yesterday, a customer asking for a discount, a banned topic. A demo is curated; production is a customer asking at ten at night about a discontinued product, three times, spelled differently.
Read the full method03
The pass criteria carry no adjectives
Each scenario defines what should happen: answer with the data, ask to narrow it down, or escalate to a person. Pass or fail. And what should happen when someone asks for a discount isn't decided by whoever built the agent — it's decided by whoever puts up the money.
Read the full method04
What a person reviews
Actions with consequences — approving, charging, promising a date — aren't closed by the agent alone: it proposes and a person signs. Where that gate sits is decided with you before we build, not after the first scare.
Read the full method05
After launch: the dashboard
An agent that's down doesn't complain, it goes quiet. So we don't just watch whether it answers: we watch how often it escalates, which questions it couldn't answer, and whether cost per conversation moved.
Read the full method06
Technical and business metrics get read together
Cost per conversation is what keeps the system honest: if it climbs while sales don't, something broke. And a lone number isn't evidence — without a before-and-after comparison, we don't publish it.
Read the full method07
A good bank is harvested, not written
The first one is written by hand: twenty or thirty scenarios drawn from how your customers actually buy. The good one is harvested later — every real conversation that went wrong becomes a new scenario, and that failure can't repeat without setting off an alarm.
Read the full method
What to ask for — here or anywhere else.
If you're evaluating vendors, ask for one thing: the scenario bank and its latest run. If they don't exist, you're buying a demo. That applies to us as much as to anyone.
What this does not promise
No natural-language system is 100% accurate, and anyone promising that is selling. What can be held to: that every fact comes from a query to your system, that anything unverifiable escalates to a person, and that when something fails there's a record to find out what happened.
Bring a real operation and we'll review it.
Fifteen minutes on your case, not on a staged demo. You leave with an honest read on what to automate first — and what to leave alone.