Our old prospecting tool made 100 claims about 20 Shopify stores. A review rewrote 91 of them. So roughly one claim in five was wrong, and the model was not the culprit. Two lines of trusted Python were. If software writes to your customers, that risk is yours.
Autonomous Technologies builds and runs the systems behind online stores. What follows is our own outreach engine, rebuilt over 418 rounds of AI agent evaluation: testing what software proves, rather than judging whether its wording sounds right. Four defects are still open, and you get to read all four.
| What was measured | Figure | Period or window |
|---|---|---|
| Wrong claims in the tool this replaced | 21%, 91 of 100 rewritten | one review pass, 20 stores, before this build |
| Builder and critic rounds logged | 418 | 7 to 19 August 2026 |
| Winners the learning loop crowned from pure noise | 35.2% of weeks, then 0.1 to 0.2% | 3,000-trial simulation, before and after |
| Components signed off, defects left open | 1 of 10, and 4 open | as at 9 September 2026 |
One claim in five was wrong, and the model was not the cause
The review traced that 21% to two lines of ordinary code, not to the writing. One read a product field that a store’s data feed never returns. The other looked for review markup that loads after the page does.
Both lines were plain Python that everyone trusted. A firmer instruction to the model would have repaired neither. Your automations carry the same shape: a stock rule reading the wrong field, a check that runs too early.
So the burden moved. Proving a claim is the software’s job, and it has to fail closed, meaning stop rather than guess, when it cannot.
The rule now in the code: the default used to be survive. A claim went out unless a word list caught it. The default is now kill. A claim dies unless the fetched page can be shown to contain it.
The builder never decides that its own work passes
We split the system into ten components and froze a written contract for each. A builder implemented one. A separate critic tried to break it, and had to return a reproduction, a failing case anyone can re-run, not an opinion. The builder fixed the cause and the piece went round again.
Every round appended one line to a log, 418 of them. The failures stayed in, and so did the rounds that started and never finished. So you can see where the work stopped.
Count the hours your team loses undoing one bad automated action. Then ask who is paid to break it before a customer meets it. That is what custom engineering and integrations is for.

A critic proved our own scoreboard was lying
One test harness, a rig that runs the real nightly job end to end, works against 25 recorded storefronts. Those are saved copies of real store pages, so a failure can be repeated. The harness kills the job at random, 20 rounds running, then checks for duplicate work. The run came back 32 of 32 green.
The critic on that harness attacked the scoreboard, not the code. A helper process was detaching itself, surviving the kill, then rewriting the evidence file. A real duplicate had been erased and the run still exited clean, 3 times out of 3.
Worse, the “no duplicate sends” check had no rows to read. It had been passing on an empty table.
Take this to your own release checks: a test that passes with nothing in it is not a test. Every check you rely on needs a guard that fails when it has nothing to count.
Two independent checks have to agree, or the claim dies
Before any sentence about a store reaches a draft, two checks must agree. They were written separately and share no code. The first confirms that every number in the claim appears in the fetched page. The second re-runs that test a different way, and makes a model quote the supporting words. A pass from the second never revives a kill from the first.
Critics attacked that part for rounds, and drove the made-up phrasings that slipped past both checks to zero on that sweep, backed by 28,525 machine-written test sentences. Then the next critic wrote a fresh set and got 153 of 1,250 through. The table below carries both halves.
What the rounds moved, and what they never measured
| What was measured | Before | After |
|---|---|---|
| Made-up phrasings passing both checks, 12-store sweep | 55 of 432, 12.7% | 0 of 432 |
| Figures of speech passing both checks | 8 of 36 | 0 of 36 |
| Escapes found by the next critic | not yet tested | 153 of 1,250, plus 410 of 500 on a new shape |
| Weeks a winner was crowned from pure noise | 35.2% | 0.1 to 0.2% |
| Replies, booked calls or revenue from the engine | not measured | not measured |
That last row is the honest one. The engine has never run unattended against a real prospect list, so no production result exists and none is invented. On the statistics: the textbook correction measured about 3%, our bar was under 1%, and the critic refused it until it reached 0.1 to 0.2%.
Human approval is structural. Campaigns are created paused, and a person activates them. The same gate runs on our marketing site, described in the SafaiKaro weekly routine.
Four defects are still open, and you get to read them
The test suite, the checks that run against the code, holds 1,826 tests across 30 files. Two error, on a function that is called and never defined. That fault has sat on the branch we ship from since 1 September 2026, inside the module this story is about.
It fails safely: the batch stops and the log shouts, rather than a wrong claim reaching an inbox. Three more defects sit open, and only 1 of the 10 components was signed off.
Here is the warning worth carrying away. This repository had no automatic test gate, so nothing re-ran the suite after the last edit. Client repositories run it on every push. Here we were the client, and we skipped it.
You should hear that from us rather than find it. We disclose the same way on client products, as in the Memox build.
Who this is for, and who it is not for
This is for you if software already acts on your behalf: pricing rules, stock syncs, refunds, order routing, anything that writes to a customer. Every wrong or repeated action costs your team hours of cleanup, and costs your name some trust.
It is not for you if the task is a one-off script a person reads first. The review loop costs more than the work. Nor is it for you if nobody on your side can say what a correct result looks like.
What we would do differently: add the test gate on day one, and give every check a guard that fails when it counts nothing.

Questions and answers
What is AI agent evaluation, in plain terms?
It is testing what software claims and does, rather than judging whether its output sounds right. Here that meant a written contract per component, a separate critic who reproduced failures, and a record that keeps the failed rounds.
What makes the critic independent?
A different role, rewarded for reproducing a failure rather than for shipping. The builder’s confidence is never the test.
How does this apply to a Shopify store?
Ask three questions of any automation you run. Does the action have evidence behind it? Does a retry repeat something the customer already saw? Can a rule change without a person approving it?
Does this prove the engine makes money?
No, and no such claim is made. There is no reply, call or revenue figure in this repository. This story is about the method and the record. Client results live in the client stories.
