On 20 March 2026, we put a star rating on a pest control website with no reviews to back it up. It stayed there for 75 days before anyone noticed, and the site was our own. SafaiKaro is a Karachi pest control business that Autonomous Technologies owns and runs, doing about ten paid jobs a month. At that size, a wrong number on the site isn't some abstract risk. It's a claim real customers read before they call, and we couldn't back it up.
| Measure | Figure | Period |
|---|---|---|
| Days an unprovable claim stayed live, first round | 75 | 20 Mar to 3 Jun 2026 |
| Days to catch the same kind of claim, third round | 0, caught and pulled the same day | 31 Aug 2026 |
| Pages carrying an unbacked trust claim | 25, down to 0 | removed 9 Sep 2026 |
| Content runs merged, first by hand then on a schedule | 4 | 6 to 21 Sep 2026 |
Nobody owned the week, so a bad claim lived until someone looked
Autonomous Technologies builds and runs online stores and local service sites. Because SafaiKaro is our own, the mistakes below are ours to report, not a client's.
Between March and August, every fix came from an audit someone happened to run. A June pass took an invented rating and a volume badge off the homepage. A July review found certification claims nobody could back up, and that fix shipped tucked inside an unrelated pricing update, because nothing required it to get its own review. Then a capability claim went live on 31 August and was pulled the same day. We only caught it because someone chose to look that week.
The other kind of miss was quieter. For 20 days in August, no changes reached the live site at all, so visitors kept seeing old copy. Nobody was watching for that gap.
We wrote down every claim we'd already gotten wrong
The fix started as a list, not a tool. A banned phrase list is a written record of claims we've gotten wrong before, kept in a config file so nobody can forget them. It now holds 18 phrases this site isn't allowed to say, and most of them are leftovers from the real removals above.
A checker reads that list every time it runs. A checker is a small script that blocks a change if it finds a problem. On new lines, it flags any banned phrase or any price that isn't in the official price list. Across the whole page, it looks for more, like broken links or a question in the page's search engine markup that never actually appears on the page.
Write your mistakes down where a machine can read them. A style guide is advice you have to remember. A versioned list of banned phrases, sitting next to the code, is the part of this system that remembers for you.
An agent proposes the change. It does not get to approve it.
Every Monday, an agent gathers the week's data and drafts a change. It opens a single pull request, which is a bundle of proposed edits waiting for a person to approve them. None of it reaches the site on its own. Two checks have to pass first, and the agent that wrote the change never runs the hostile review.
The first check is a lint gate. That's a rule built into the process itself, so nobody has to remember to run a separate audit. It's the same checker from above, just applied to this week's change.
The second check is a hostile review. A fresh agent, one that never saw how the drafting agent reasoned, has a single job: find a reason to reject the change. It judges against named rules for claims, price, local geography and tone. The change gets one chance to be fixed. If it fails a second time, the piece is dropped, and the pull request records why. We use the same drafter and critic split throughout our outbound engine build.
Only once both checks pass does the pull request wait for a person, and it never merges itself. Our founder approves every merge, and nobody is named to cover for him. That's the thinnest part of the system, and it's the first thing on our fix list below.
Four weekly runs have merged, and one caught something the checker couldn't
We started the first run by hand. Since then, three more have opened on their own, right on schedule, though each moved at a very different pace.
| Run | Opened | Time to merge | What made it different |
|---|---|---|---|
| First | By hand, Sun 6 Sep 2026 | 17 minutes | The only one not started by the schedule |
| Second | On schedule, Mon 7 Sep, 08:30 Karachi | 16 minutes | First run nobody had to start |
| Third | On schedule, Mon 14 Sep, 08:26 Karachi | 3 days | A leftover termite-guarantee line both checks missed, fixed at the merge review |
| Fourth | On schedule, Mon 21 Sep, 08:25 Karachi | 8 minutes | Fixed a namesake-company mix-up, see below |
The fourth run shipped a fix for something the lint gate has no way of checking. Once a month, the routine runs two independent web searches to see what AI answer engines are saying about SafaiKaro. The September round found them describing a different company abroad with a similar name. So we added identity lines to the site's footer and its machine readable summary, which gives next month's check something to confirm against.
The gate is a sentence in a prompt, not a setting in GitHub. Searching the workflow files for the checker returns nothing. No required check runs it. A run that skipped it would open a pull request that looks just like one that passed.
What we would do differently, and who this is not for
Three fixes are still open:
- Make the checker a required check, so the gate is more than just an instruction.
- Tighten the price rule, so a cost figure can't slip through as a real price.
- Name a second approver. Right now, with only one, every ready pull request waits on one person's calendar.
We're not claiming any lead, sale or ranking numbers here. The measured lift on this site came from a change a person shipped six weeks earlier, and that credit belongs to the local search work, not to this routine.
This is for you if your store makes promises in its copy, your prices show up on more than one page, and you're the only person who'd catch a wrong one. It's not for you if you want a number to move this quarter, or if you're after an agent that publishes with nobody standing in its way.

Questions and answers
Can an AI content workflow publish on its own?
No. It drafts a branch and opens one pull request. A person has to merge it by hand. GitHub does not enforce that limit yet, which is the first item on our fix list.
What makes an AI content workflow safe to run on a live site?
Four things, in order: collect evidence first, check the draft against a banned-phrase list, have a second agent review it without seeing the first agent’s reasoning, and let only a person merge.
Does the checker stop every false claim?
No, it catches named phrases and broken structure only. A namesake mix-up on an AI answer engine needed its own monthly check, and one wording error got past both checks and was caught by a person at merge.
What should land on your desk each week?
One proposed change, why it was proposed, and what it passed. As of 23 September 2026, 7 of 54 queued items were dropped, one after a failed second review, and 4 were handed to a person.
How does this apply to an ecommerce store?
The order stays the same and the artifacts change. Price checks read your price rules, deploy health reads your theme version, and the pages that matter are your collections and product pages.
