Skip to main content

AI workflow engineering · Autonomous Technologies

AI agent evaluation: 418 critic rounds, and 4 defects we published

One in five claims our old prospecting tool made was wrong. We rebuilt it so a separate critic had to break every part first, and we published the defects it left open.

Autonomous website introducing Shopify and ecommerce engineering
Autonomous's public website, captured 11 September 2026. The engine here is our own tooling, so every defect below is ours to publish.
Integrations & Custom Systems6 minute readBy Rizwan QaiserRead the case study

Our old prospecting tool made 100 claims about 20 Shopify stores. A review rewrote 91 of them. So roughly one claim in five was wrong, and the model was not the culprit. Two lines of trusted Python were. If software writes to your customers, that risk is yours.

Autonomous Technologies builds and runs the systems behind online stores. What follows is our own outreach engine, rebuilt over 418 rounds of AI agent evaluation: testing what software proves, rather than judging whether its wording sounds right. Four defects are still open, and you get to read all four.

Where this build started, 21% wrong claims, and where it ended, 0.1 to 0.2% false winners
What was measuredFigurePeriod or window
Wrong claims in the tool this replaced21%, 91 of 100 rewrittenone review pass, 20 stores, before this build
Builder and critic rounds logged4187 to 19 August 2026
Winners the learning loop crowned from pure noise35.2% of weeks, then 0.1 to 0.2%3,000-trial simulation, before and after
Components signed off, defects left open1 of 10, and 4 openas at 9 September 2026
Where this build started, 21% wrong claims, and where it ended, 0.1 to 0.2% false winners

One claim in five was wrong, and the model was not the cause

The review traced that 21% to two lines of ordinary code, not to the writing. One read a product field that a store’s data feed never returns. The other looked for review markup that loads after the page does.

Both lines were plain Python that everyone trusted. A firmer instruction to the model would have repaired neither. Your automations carry the same shape: a stock rule reading the wrong field, a check that runs too early.

So the burden moved. Proving a claim is the software’s job, and it has to fail closed, meaning stop rather than guess, when it cannot.

The rule now in the code: the default used to be survive. A claim went out unless a word list caught it. The default is now kill. A claim dies unless the fetched page can be shown to contain it.

The builder never decides that its own work passes

We split the system into ten components and froze a written contract for each. A builder implemented one. A separate critic tried to break it, and had to return a reproduction, a failing case anyone can re-run, not an opinion. The builder fixed the cause and the piece went round again.

Every round appended one line to a log, 418 of them. The failures stayed in, and so did the rounds that started and never finished. So you can see where the work stopped.

Count the hours your team loses undoing one bad automated action. Then ask who is paid to break it before a customer meets it. That is what custom engineering and integrations is for.

How the system works: Written contract, Builder implementation, Adversarial test, Reviewed correction
Four steps, repeated until the attacks stop landing. The critic reproduces the failure, and the fix goes to the cause.

A critic proved our own scoreboard was lying

One test harness, a rig that runs the real nightly job end to end, works against 25 recorded storefronts. Those are saved copies of real store pages, so a failure can be repeated. The harness kills the job at random, 20 rounds running, then checks for duplicate work. The run came back 32 of 32 green.

The critic on that harness attacked the scoreboard, not the code. A helper process was detaching itself, surviving the kill, then rewriting the evidence file. A real duplicate had been erased and the run still exited clean, 3 times out of 3.

Worse, the “no duplicate sends” check had no rows to read. It had been passing on an empty table.

Take this to your own release checks: a test that passes with nothing in it is not a test. Every check you rely on needs a guard that fails when it has nothing to count.

Two independent checks have to agree, or the claim dies

Before any sentence about a store reaches a draft, two checks must agree. They were written separately and share no code. The first confirms that every number in the claim appears in the fetched page. The second re-runs that test a different way, and makes a model quote the supporting words. A pass from the second never revives a kill from the first.

Critics attacked that part for rounds, and drove the made-up phrasings that slipped past both checks to zero on that sweep, backed by 28,525 machine-written test sentences. Then the next critic wrote a fresh set and got 153 of 1,250 through. The table below carries both halves.

What the rounds moved, and what they never measured

Before and after, measured by the critics during the build, none of it audited externally
What was measuredBeforeAfter
Made-up phrasings passing both checks, 12-store sweep55 of 432, 12.7%0 of 432
Figures of speech passing both checks8 of 360 of 36
Escapes found by the next criticnot yet tested153 of 1,250, plus 410 of 500 on a new shape
Weeks a winner was crowned from pure noise35.2%0.1 to 0.2%
Replies, booked calls or revenue from the enginenot measurednot measured
Before and after, measured by the critics during the build, none of it audited externally

That last row is the honest one. The engine has never run unattended against a real prospect list, so no production result exists and none is invented. On the statistics: the textbook correction measured about 3%, our bar was under 1%, and the critic refused it until it reached 0.1 to 0.2%.

Human approval is structural. Campaigns are created paused, and a person activates them. The same gate runs on our marketing site, described in the SafaiKaro weekly routine.

Four defects are still open, and you get to read them

The test suite, the checks that run against the code, holds 1,826 tests across 30 files. Two error, on a function that is called and never defined. That fault has sat on the branch we ship from since 1 September 2026, inside the module this story is about.

It fails safely: the batch stops and the log shouts, rather than a wrong claim reaching an inbox. Three more defects sit open, and only 1 of the 10 components was signed off.

Here is the warning worth carrying away. This repository had no automatic test gate, so nothing re-ran the suite after the last edit. Client repositories run it on every push. Here we were the client, and we skipped it.

You should hear that from us rather than find it. We disclose the same way on client products, as in the Memox build.

Who this is for, and who it is not for

This is for you if software already acts on your behalf: pricing rules, stock syncs, refunds, order routing, anything that writes to a customer. Every wrong or repeated action costs your team hours of cleanup, and costs your name some trust.

It is not for you if the task is a one-off script a person reads first. The review loop costs more than the work. Nor is it for you if nobody on your side can say what a correct result looks like.

What we would do differently: add the test gate on day one, and give every check a guard that fails when it counts nothing.

Rizwan, founder of Autonomous Technologies and the author of this case-study collection.
Rizwan, founder of Autonomous Technologies, supervised all 418 rounds and published the four defects left open.

Questions and answers

What is AI agent evaluation, in plain terms?

It is testing what software claims and does, rather than judging whether its output sounds right. Here that meant a written contract per component, a separate critic who reproduced failures, and a record that keeps the failed rounds.

What makes the critic independent?

A different role, rewarded for reproducing a failure rather than for shipping. The builder’s confidence is never the test.

How does this apply to a Shopify store?

Ask three questions of any automation you run. Does the action have evidence behind it? Does a retry repeat something the customer already saw? Can a rule change without a person approving it?

Does this prove the engine makes money?

No, and no such claim is made. There is no reply, call or revenue figure in this repository. This story is about the method and the record. Client results live in the client stories.