A customer emails at 11pm to change a delivery address. One tool drafts the reply your team sends in the morning. Another cancels the fulfilment, rewrites the address in the system that holds your orders, and reissues the label before anyone wakes up. Both get sold to you as AI. Only the second one can cost you a remade order by breakfast.
Generative AI produces content: a draft, an image, a block of code. A person then decides what happens to it. Agentic AI decides and acts: it picks the next step, writes to a live system, and carries on without asking. Autonomous Technologies builds and runs the systems behind Shopify and Shopify Plus stores, so operators ask us this most weeks.
Pick generative AI when a person still reads every output before it leaves: product copy, support replies, a first pass at a policy. Pick agentic AI when the cost is the re-keying rather than the writing, and only once an undo path exists. The split is not intelligence. It is who owns the failure.
| The question you are answering | Generative AI | Agentic AI |
|---|---|---|
| What leaves the system | 1 draft, still inside your building | 1 or more writes into a live record |
| Who signs off on each output | A person, every time | A person once at design time, then nobody |
| Blast radius of one wrong output | 1 message or 1 page | Every record that run touched |
| What must exist before you switch it on | A prompt and a reviewer | A tested undo, an audit log and a stop button |
| What broke it in our own build | The model's wording: 0 of 2 root causes | The action code: 2 of 2, behind a 21% wrong-claim rate |
| Where the projects die | Not measured anywhere we could find | Over 40% canceled by end of 2027, Gartner, June 2025 |
This is for the operator accountable for a Shopify or Shopify Plus store after launch: orders re-keyed into an ERP, stock drifting between systems, pricing rules split across four apps. An ERP here is the system of record holding your stock, costs and orders. This is not for you if nobody on your side can state what a correct result looks like. Without that sentence, neither tool is safe, and an agent is worse than a draft.
Generative AI is cheap to be wrong with, because a person reads it first
Generative AI is a model that produces new content from a prompt. That content is text, an image, code or a summary. It does not touch your systems. Its output lands in a draft, a document or a chat window, and a person decides whether it goes anywhere.
That makes ownership simple. The reviewer owns the failure. If the model invents a fabric composition, your merchandiser catches it before it reaches a product page.
The running cost is that reviewer’s attention, and it rises with volume. Every draft is a read. You are buying a faster first version, not fewer people.
Good fits in a store: descriptions for a new range, first-pass replies in a support queue, alt text for a catalogue backlog, a plain summary of last month’s return reasons. Bad fit: anything nobody is going to read before it ships.
Agentic AI acts, so the undo path is the part you actually build
Agentic AI is a system where the model chooses the next step and then takes it. The tools it uses write to real systems. Anthropic’s engineering team draws the line at who directs the process. In a workflow, code decides the path. In an agent, the model decides it, and the same page warns of “higher costs, and the potential for compounding errors” (“Building effective agents”, 19 December 2024).
The model is the same model. What changes is that a wrong step now writes to your data. Every hour your team spends unwinding those writes is the real bill. Cutting that bill is the whole point of custom engineering and integrations, and it is scaffolding work, not model work.
Three things have to exist before an agent touches anything live. A tested undo, meaning a proven way to reverse every write a run made. An audit log, meaning a record of what it did and why, in words a person can read. A stop button that halts a run mid-flight without leaving half-finished work behind.
Good fits in a store: the handoffs nobody enjoys. Orders re-keyed from Shopify into an ERP or a 3PL, which is a third-party warehouse that picks and ships for you. Stock counts drifting apart in two systems. Refunds that need four screens. Bad fit: any write that is expensive to reverse and that nobody can check. That choice between a fixed rule and a deciding agent is its own decision, worked through in our guide to workflows or agents.
The failure worth comparing is the boring one, not the clever one
Most published comparisons argue about reasoning. Operators lose money somewhere duller. A field comes back empty, a check runs before the page finishes loading, a job dies halfway and starts again.
Generative AI meets those faults in front of a reader. Agentic AI meets them inside your records.
| The fault | With generative AI | With agentic AI |
|---|---|---|
| The model states something untrue | The reviewer deletes the line, 1 draft affected | It is already written, and the run keeps going |
| A field it reads is always empty | The output reads thin and somebody notices | Every record in the run is written from nothing |
| The run dies halfway through | You lose 1 draft | You get duplicates, unless the code was built to prevent them |
| A safety check has nothing to check | You read the output anyway | Ours reported 32 of 32 green with a real duplicate erased |
In our own agent the model’s wording caused none of the wrong claims
We rebuilt our own outreach engine over 418 builder and critic rounds, 7 to 19 August 2026. The tool it replaced had made 100 claims about 20 Shopify storefronts. A review rewrote 91 of them, so roughly one claim in five was wrong.
The root cause was not the writing. It was two lines of ordinary Python. One read a product field that a store’s data feed never returns. The other looked for review markup that loads after the page does. No change to a prompt would have fixed either.
The rule that came out of it, taken from the code comment: the default used to be survive. A claim went out unless a word list caught it. The default is now kill. A claim dies unless the fetched page can be shown to contain it.
| What we measured | Before | After |
|---|---|---|
| Made-up phrasings passing both evidence checks | 55 of 432, 12.7% | 0 of 432 |
| Escapes found by the next critic, fresh test set | not yet tested | 153 of 1,250 |
| Weeks a winner was crowned from pure noise, 3,000-trial simulation | 35.2% | 0.1 to 0.2% |
| Components signed off, defects left open | 0 of 10 signed off | 1 of 10, 4 defects open on 9 September 2026 |
| Replies, booked calls or revenue from the engine | not measured | not measured |
That last row is the honest one. The engine has never run unattended against a real prospect list, so no commercial result exists and none is invented here.
We let an agent act only once the undo path has been tested
Our harness kills the running job at random, 20 rounds against 25 recorded storefronts. It then checks for duplicate work. The run came back 32 of 32 green. A critic proved the scoreboard was lying. A helper process survived the kill and rewrote the evidence file, so a genuine duplicate had been erased, 3 times out of 3.
The check for duplicate sends was reading a table with 0 rows in it. It had been passing on nothing.
Carry this into your own release checks. A test that passes with nothing in it is not a test. Every check you rely on needs a guard that fails when it has nothing to count.
So the order we use is fixed. Rules first, then a generative draft with a reviewer, then an agent, and only where re-keying is the actual cost. The engineering side of that order is set out in our guide to agentic systems. Campaigns in our own engine are still created paused, and a person activates them.
The same gate holds on client work. On Sene Studio’s order systems, 37 changes reached the branch behind real garment orders between 14 July and 8 September 2026. A coding agent wrote 35 of them. A named person accepted every one, and 5 were sent back before they could ship. A defect caught at review costs minutes. The same defect caught after a factory run costs a garment.
Agent projects die on cost and control, not on the model
Gartner’s June 2025 prediction: over 40% of agentic AI projects canceled by the end of 2027. The reasons it gave were escalating cost, unclear business value and inadequate risk controls. Its analysts also counted roughly 130 vendors with genuine agentic capability out of thousands claiming it, as reported on 29 April 2026.
Read that as a buying warning rather than a verdict on the technology. Most of what is sold as an agent is a rules engine with a model bolted on the front.
You may not need either tool. Shopify Flow watches your store for an event, checks a condition and performs an action. No model decides anything, so nothing can drift. If a rule covers the job, a rule is the cheaper thing to own.
The other break is a store that changes faster than its scaffolding. An agent built against last season’s fulfilment logic will keep acting confidently after the logic moves. Somebody has to own that drift, by name. Our own count says the same: 1 of 10 components signed off and 4 defects still open on 9 September 2026, on software we wrote ourselves.
Agentic AI and generative AI, answered
What is the difference between agentic AI and generative AI, in one sentence?
Generative AI produces content that a person decides what to do with. Agentic AI decides and then acts on a live system, so the decision and its consequence both land before anyone reads them.
Is an agent just a generative model with tools attached?
Largely, and that is the point. Anthropic’s “Building effective agents”, 19 December 2024, puts the line at who directs the process: code in a workflow, the model in an agent. The model does not get smarter. Your data gets reachable.
Where should a Shopify operator start?
With the handoff that costs the most hours, usually orders re-keyed from Shopify into an ERP or a 3PL. Try deterministic rules first. Shopify Flow already fires on an event and performs an action without a model in the path.
Why do so many agent projects get cancelled?
Gartner’s June 2025 prediction puts it at over 40% by the end of 2027, on cost, unclear value and weak risk controls. In our experience the cost lands in the scaffolding: the undo, the log and the checks that have to fail loudly.
Can we run both together?
That is the usual answer, and it needs a person at the end. An agent selects the 20 accounts that need a follow-up and writes the record. A generative model drafts each message. Your team sends them.
Start with the one handoff that costs you the most hours to undo
Count the hours your team spends undoing one bad automated write. A duplicate refund. An order that moved twice. Stock that never came back. Bring us that one handoff.
You leave with the failure cases anything acting on it has to survive. You also leave with a plain answer on whether a rule, a draft or an agent is the right shape for it.
The fear of starting an AI project that gets quietly cancelled is the right fear to have. It is why the first conversation produces a map of your systems rather than a build.



