Your support lead pasted 40 order questions into a chatbot last Tuesday and cleared them in an hour. Nobody told you. Two answers quoted a shipping window you stopped offering in June. That is the first LLM decision to reach a CEO, and it has nothing to do with which model you buy.
A CEO owns four LLM decisions: which job the model touches, who checks its output before a customer sees it, how you count the time it gives back, and what turns it off. An LLM, or large language model, is software that predicts text from text. It is fast, confident, and sometimes wrong.
This guide is for the operator accountable after launch
You run a Shopify or Shopify Plus store. Orders land in an ERP. Stock counts drift between systems. Pricing rules live in three apps, and one person holds the map in their head. Now somebody wants to put a model in the middle of that.
This is not for the AI lead picking a vendor. It is not for the engineer tuning prompts. Both jobs are real, and both sit downstream of the four decisions here. The question a CEO answers first is who owns a wrong answer.
That question is the same one a systems map answers. Our engineering work starts by naming every system that touches an order and the person who owns each break. A model is just one more system that touches the order.
The four decisions sit with you, and each carries a number
Nobody else in the building can make these calls. Your head of support cannot decide what the company will disclose. Your developer cannot decide what an error costs. Each decision has one number attached, and none of the four is a model benchmark.
| Decision | What you decide | The number you check | When you check it |
|---|---|---|---|
| 1. The job | Which task the model touches first | Hours per week on that task, measured over 4 weeks | Baseline 4 weeks before, 4 weeks after |
| 2. The checker | Who reads the output before a customer does | Share of outputs edited before they ship, per 100 sent | Weekly, for the first 8 weeks |
| 3. The return | Whether the time actually comes back | Orders per week still re-keyed by hand, against your week-1 count | Monthly |
| 4. The exit | What switches it off, and who can do it | Hours from a bad answer to shutdown, target under 24 | Per incident |
Decision one: name the job before you name the model
Most LLM projects start at the wrong end. Someone picks a tool, then hunts for work to give it. Start with a task you can already describe in one sentence, that happens weekly, and that somebody on your payroll currently does by hand.
Good first jobs in a store look boring. Draft a reply to an order-status email. Summarise the week’s exceptions report. Write a first pass of a product description. Each one produces text a human reads before anything ships.
Bad first jobs look exciting. Answering customers live, setting prices, deciding refunds. Those put an unreviewed output in front of a customer or a ledger on day one.
The rule: pick a job where a wrong answer is visible and cheap to fix. Your first LLM job should fail in front of an employee, never in front of a customer.
Decision two: a tool that fails one task in three needs a named checker
Stanford HAI’s 2026 AI Index Report records AI agents moving from 12% to about 66% task success on OSWorld, a benchmark of real computer tasks. That is a large jump in one year. It also means agents still fail roughly one attempt in three on structured benchmarks. The same report counts 362 documented AI incidents in 2025, up from 233 in 2024.
So the checker is not a nice-to-have. It is the operating cost of the tool. Name a person, put the work in their week, and count what they change. When the edit rate stops falling, you have learned where the model is reliable.
Autonomous connects and runs the systems behind Shopify and Shopify Plus stores: the integrations to an ERP, a 3PL or accounting, the automation between them, and the pricing and merchandising rules. We also build and run our own AI agent product, Memox, so this is not theory for us. In August 2025 we replaced a third-party no-code flow tool with an agent layer we control, because we needed our own hands on the agent’s tools, state and tracing. Human handover is a tool the agent can call, not a bolt-on. The Memox case study sets out what that ownership changed.
| Fact | Figure | Period | What it changes for you |
|---|---|---|---|
| Agent task success on OSWorld | 12% to about 66% | 2025, per the 2026 AI Index | Budget for checking: 1 in 3 structured tasks still fails |
| Documented AI incidents | 362, up from 233 | 2025 against 2024 | Incident reporting is growing faster than safety reporting |
| Organisational AI adoption | 88% | 2025 | Your staff are already using these tools, with or without a policy |
| EU AI Act transparency rules | In effect August 2026 | Act applicable 2 August 2026 | Selling into the EU means telling people when they talk to a machine |
Sources: the 2026 AI Index Report from Stanford HAI, read 21 September 2026, and the European Commission’s AI Act page, read the same day.
Decision three: count the hours returned, not the tools bought
The usual pitch is a saving. Ten hours a week, five people freed up, a support queue cut in half. Ask one question in return: what was the number before?
Most companies cannot answer it. That is the whole problem. Count the task for four weeks before the model touches it. Count the same task for four weeks after. Keep the two counts in the same units, done by the same person, on the same day of the week.
If the hours come back but land somewhere nobody tracks, you have moved work, not removed it. That shows up fast in a store: the support queue shrinks, and the finance team quietly starts re-keying the orders support used to catch.
Warning: an hours-saved figure with no baseline is a story, not a result. If nobody counted the four weeks before, you do not have a saving. You have a feeling.
Decision four: the disclosure dates are already set
The European Commission’s AI Act page states that transparency rules come into effect in August 2026, and that people must be told when they are dealing with a machine rather than a person. Providers of generative AI also have to make AI-generated content identifiable.
The wider timeline is fixed. The Act entered into force on 1 August 2024. It became applicable on 2 August 2026. Obligations for general-purpose AI models applied from 2 August 2025.
You do not need a legal opinion to act on that. You need a line on the page, a record of what the model was allowed to say, and a person who can answer when a regulator or a customer asks.
The rule: if a customer can read it, label it. A sentence saying an assistant is automated costs you nothing and settles the question before it is asked.
What breaks, and who owns it
Four failures show up again and again. Each has an owner, and the owner is never the vendor.
The stale answer. Your returns window changes and the model keeps quoting the old one, because nobody updated the documents it reads. Owner: whoever owns the policy, not whoever owns the tool.
The silent pilot. Three teams run three different tools on three different cards, and nobody counts anything. Owner: you, because only you see all three.
The confident mistake. The output is fluent, specific and wrong, so it survives a skim read. Owner: the named checker from decision two.
The leak. Customer data, supplier prices or draft contracts go into a tool nobody reviewed. Owner: whoever signed the agreement.
The objection we hear most is that a store’s data is too specific for a model to be useful. That is usually right about the final answer and wrong about the draft. Retrieval over your own documents changes what the model can see, and it is cheaper than training anything.
All four decisions start from the same count, and most stores do not have it. How many hours a week go into re-keying orders between Shopify and the systems behind it? How many orders get touched twice before they ship? Those hours are the budget you would spend on checking a model’s work. Find them first.
Questions CEOs ask about LLMs
What is a large language model, in one sentence?
A large language model is software trained on text that predicts the next piece of text, which is why it can draft, summarise and translate, and why it can be fluently wrong.
Do I need to train a model on my own data?
Almost never at the start. Retrieval over your existing documents gives the model your policies and product facts without training anything, and it is far easier to correct when a policy changes.
Should an automated assistant answer customers on my store today?
Only with two things in place: a visible line saying it is automated, and a path to a human. The EU AI Act transparency rules come into effect in August 2026, so the disclosure is not optional if you sell there.
Who should own this work, IT or marketing?
Whoever owns the process the model touches. If it drafts support replies, support owns it. IT owns access and data, and you own the exit.
How do I know a pilot is working?
The edit rate falls week on week, and the hours you counted before the pilot are lower after it. If neither number exists, the pilot has not started.



