On 13 May 2026, we sent a merchant a Shopify store audit claiming their best product pages were broken. They weren't. Those pages showed backorder wording next to an Add to Cart button that worked just fine, and shoppers could buy without any trouble. Our mistake was simple: we'd searched the page code for the word "Unavailable," found it, and called it a problem.
If you've ever had a report tell you something about your store that one click proves wrong, you know what it costs. Your developers lose a week chasing a problem you never had, while the real leak keeps draining money in the background.
| What changed | Figure | Period |
|---|---|---|
| How far the wrong finding was overstated | about 4 to 7 times, HIGH cut to MEDIUM | audit sent 13 May, revalued 26 May 2026 |
| Days from the bad audit to the first fix | 13 | 13 to 26 May 2026 |
| Red team reviews before a customer report | 4, three of them before any dollar figure | enforced in code since 31 July 2026 |
| Store audits with a finished customer report, counted from disk | 15 | 13 July to 27 August 2026 |
Hazn is the system every audit and deliverable runs through
Hazn is the internal delivery system behind Autonomous Technologies, a Shopify systems agency. Its plugin manifest lists 16 plugins, each built for a single job, like the Shopify store audit, prospect research, lean store scans, proposals and case studies. What that means for you is simple: a lesson we learn on one store protects the next one.
That's why one wrong finding ended up changing the whole agency. The fix didn't stay inside a single report. It went into the shared review that every audit loads, where the May case now sits as the example of what not to do.
Rizwan, our founder, wrote 200 of the 229 commits up to 20 September 2026. Abdullah, our cofounder, wrote the May 2026 decision record that split Hazn into plugins. Two more commits came from Haaris, an AI agent the founder runs.
Five buying states decide which finding we may write
A buying state is simply what a shopper can do on a product page right now. Before anyone writes a single finding, every page we sample gets sorted into one of five states.
| Buying state | Can a shopper buy? | Inventory finding allowed? |
|---|---|---|
| In stock | Yes | No |
| Backorder, still purchasable | Yes | No, a wording note at most |
| Pre-order | Yes | No, a lead-time note at most |
| Sold out with a notify form | No | Yes |
| No buy button at all | No | Yes |
Every finding has to point back to three things: the capture it came from, the check that ran, and the saved state of the page. If any of those is missing, the page goes back for another capture instead of into your report.
Four red teams argue with every figure, and three of them go first
A red team is a review whose only job is to attack the work. Hazn runs four of them, in an order that's fixed in code. The first two run side by side. One asks whether the starting sales figure is real, and the other asks whether a store operator would actually agree with the finding. The third checks how big each claimed gain really is. Only after all three have had their say does any dollar figure get produced. Then the fourth reads the result the way your finance lead would.

The May failure wasn't an arithmetic problem. The maths was fine. It was the meaning that was wrong. Nothing in the process asked the question any operator would ask in three seconds: can I buy this right now? So in June 2026 we added a live recheck, and since July it runs first. Before the red teams see anything, every claim behind a figure gets checked again on the live store.
Written rules did not hold, so the gates moved into code
A written rule can't stop a bad report. A gate in code can. Our first fix, on 26 May 2026, was five written rules in a single commit, and not one of them could actually stop a run. Code that could halt an audit didn't arrive until July 2026.
On 31 July 2026, the whole audit moved into one program with 23 stages that always run in the same order. We confirmed that order by reading it straight from the code on 23 September 2026. No stage can start until every stage before it has passed. The revenue model alone has seven hard gates, and any one of them can block the report. And if there's no trusted starting sales figure, it won't print a single dollar figure.
A written rule and a gate in code are different controls. A rule asks. A gate stops the run. When a vendor shows you a process document, ask which parts of it can halt the work.
| What the audit does | 13 May 2026 | 23 September 2026 |
|---|---|---|
| Evidence behind an inventory claim | a text match in theme code | a browser capture, its page state and a check ID |
| Buying states separated | 0 | 5, and 2 allow an inventory finding |
| Review order | by instruction | 23 stages enforced in code |
| Gates that can halt a report | 0 | 7 in the revenue model |
| Automated tests in the audit pipeline | not counted | 290 passed |
| Findings challenged by merchants | not measured | not measured |
We didn't delete the failed audit. It lives on in the system as a labelled test case, so anyone who edits the rules runs into it first.
What we would do differently, and who this is not for
We'd never again ship a set of written rules and call it a fix. Of the five rules from May, one is gone and two have moved into other files.
We also don't have a build server. Our 290 tests are run by hand. The last time we ran them, on 23 September 2026, they all passed. Every gate runs inside the audit itself. We also haven't counted how many merchants have challenged a finding since May, and zero out of a number nobody counted isn't a result.
What you get back is time, because your team fixes the leaks that are actually real, first. This is for you if you run a Shopify store and keep getting handed figures to act on. It's not for you if you want a scored checklist by Friday. This audit takes longer, and it'll say "unknown" where a cheaper report would hand you a confident zero.
Questions and answers
Why is a rendered capture needed when the page code is right there?
Theme code carries every stock message at once, including ones no shopper sees. Only a capture of the rendered page shows whether the buy button works and which message is on screen.
What is a red team in a Shopify store audit?
A red team is a review whose only job is to attack the work. In an Autonomous Technologies audit, four red teams check the starting sales figure, the operator’s reading, the size of each gain and the finance case.
What happens when one of your findings turns out to be wrong?
At Autonomous Technologies, the correction goes into the evidence or the model input first, then the report. The May 2026 case now sits in the shared review as the example every audit meets.
Are the revenue figures in an audit guaranteed?
No. They are estimates built on evidence plus stated assumptions, and a good report separates a proven defect from an uncertain impact.
