Skip to main content

RAG pipelines for AI: architecture, evaluation and production controls

Build a RAG pipeline for grounded AI answers. Plan ingestion, access filtering, citations, source updates and evaluation before moving beyond a prototype.

Rizwan QaiserOctober 6, 20269 min read
Discuss RAG pipelines
An editorial illustration of permitted evidence moving from an approved source to a grounded answer.
Retrieve permitted evidence. Answer with a source.

Three takeaways

  • RAG has connected indexing and answering paths.
  • Retrieval must preserve access and freshness rules.
  • Evaluate retrieval and citations before you scale.

A support agent follows the assistant and refunds an order that should have been exchanged. The procedure it quoted was withdrawn in March. Nobody removed it from the index, so it is still the closest match. The assistant broke none of its own rules, and that is the problem.

The short answer. A RAG pipeline finds approved source material for a question and supplies it to a language model so the answer is grounded in that material. Three things decide whether it holds up in production: the architecture, meaning the indexing and answer paths; the evaluation set somebody keeps; and the production controls, meaning access filtering, citations and a set way to act when evidence is weak. An embedding model and a chat box are not a pipeline.

RAG stands for retrieval-augmented generation. Retrieval means finding the records that fit. Grounded means the answer is held to what was found. The original RAG paper framed it as a model’s own knowledge plus fetched outside memory. In a company the question is blunter. Can the user see which approved record backs the answer? Can the system say it does not know?

One wrong answer costs a refund, a reopened ticket and an hour of a manager’s day. Autonomous Technologies is an engineering firm that builds and runs the data and AI systems behind Shopify stores and commerce businesses. Our clients are founders and operators who answer for what a customer is told. Our engineering work on assistants starts with the source boundary. Most bad answers we are asked to explain trace back to a document that should not have been in the index.

This is for whoever answers for the wrong answer

Read this if you are building or buying a knowledge assistant, support copilot or internal search over company content. It is for whoever has to explain a wrong answer to a customer. It assumes the content exists and somebody publishes it.

Skip this if the procedures are not written down, or nobody owns them. Retrieval will surface that mess at speed. Fix the source of truth first.

What is a RAG pipeline?

A RAG pipeline has two paths. The indexing path turns approved source material into records you can retrieve. The answer path takes a question, fetches only the records that user may see, picks the usable evidence, and asks a model for a bounded answer.

The promise is traceability, not omniscience. If a policy changed, the pipeline should surface the current approved version and name its source.

Stage
Source intake
Responsibility
Accept approved content. Record owner, version, access policy.
Evidence to retain
Source id, change time, owner, permissions.
Stage
Prep and indexing
Responsibility
Parse, cut up, enrich, index. Keep provenance.
Evidence to retain
Parser version, chunk map, index version.
Stage
Retrieval
Responsibility
Find candidate evidence for this user and question.
Evidence to retain
Query, source ids, filters, scores.
Stage
Response
Responsibility
Bound the answer by the evidence and the policy.
Evidence to retain
Model and instruction version, citations, state.
Stage
Evaluation and operations
Responsibility
Test behavior. Catch regressions. Improve under change control.
Evidence to retain
Test run, incident record, owner, release call.
The minimum responsibilities in a production RAG pipeline.

Start with a source and user boundary

The first decision is what the assistant may know. Start with one narrow collection that has an owner and a change process. Published support articles will do. Do not index every shared drive. Broad intake mixes stale, private and clashing material before you can see how retrieval behaves.

Permissions travel with the content. A user who cannot open a document at source must not get it from the assistant. That is a retrieval rule, not a filter bolted on after the answer.

The Google Cloud RAG Engine overview sets out a managed flow in five steps: ingest, clean and chunk, index, retrieve, generate. A managed service still does not replace your source ownership.

Approved documents move through parsing, segmented records and metadata enrichment into an index, while source ownership, access policy and freshness checks govern each step.
Illustrative indexing controls. A retrievable passage should retain a link to its current approved source and permission context.

Build the indexing path for evidence, not just similarity

Preparation shapes what the system can later cite. Keep the title, source link, section heading, version, owner, access labels and a stable pointer to the original. Cut material at real boundaries where you can: headings, steps or product sections. A chunk is one retrievable piece of a document. Cutting chunks at a fixed length can split a rule from the condition that makes it safe.

Metadata is the set of labels on each chunk, such as owner, version and market. Metadata makes filters possible. A policy assistant may filter by country, unit, product version and published state. A support assistant may filter by plan and language. Store only filters an owner can fill in and keep true. Half-filled metadata creates false confidence, because it drops the right source in silence.

Freshness needs a defined behavior. Pick an update signal, an indexing schedule, and a visible state for content not yet indexed. If a source is withdrawn, the pipeline must stop retrieving it.

Make the answer path selective and explainable

The answer path is a run of decisions. Read the question. Apply the permission and scope filters. Fetch candidates. Rerank them, which means re-ordering the candidates with a closer read. Judge whether the evidence is enough. Then answer with citations, or fall back safely.

Illustrative retrieval path

Indexing preserves evidence; the query path decides whether that evidence is usable for this person and question.

Indexing lane

Approved source
Parse at meaningful boundaries
Attach owner, version and access metadata
Versioned index

Query retrieval reads this index →

Query lane

User question
Apply identity and scope filters
Retrieve and rerank allowed passages
Cited answer or safe fallback

Do not use a low similarity score as a blanket refusal threshold. Score ranges differ by model, index and corpus. Test real question types instead. Then define what the product does in each failure state: no source, clashing sources, stale source, blocked source, and a vague question.

Make the evidence legible. Name the source so the user can judge it. A citation is not proof the answer is right. It gives the user a way to check.

Grounded means the answer is constrained by retrieved evidence. It does not mean the source is current, complete or right for every user.

Evaluate RAG as a system, not a chat demo

Judge the whole pipeline, from source record to the answer a user reads. Check four things on each example: the right source came back, the citation points at the text, the answer stays inside the evidence, and the fallback fired when it should. An answer can sound fluent and cite the wrong document. It can cite the right one and add a step that is not in it. It can find the right text and show it to the wrong person. Each of those needs a different fix.

Question type
Clear answer in one approved source.
Of a first 30
8
A pass looks like
right answer, citation on the section.
Question type
Answer needs a version or market filter.
Of a first 30
6
A pass looks like
the filter runs before retrieval.
Question type
Two approved sources disagree.
Of a first 30
4
A pass looks like
it names the clash, it does not merge it.
Question type
Source is withdrawn or stale.
Of a first 30
4
A pass looks like
the record is gone, or the version is shown.
Question type
User may not see the source.
Of a first 30
4
A pass looks like
nothing comes back, nothing is summed up.
Question type
Hostile text inside a fetched document.
Of a first 30
4
A pass looks like
it is ignored, and the attempt is logged.
A first review set of 30 questions, weighted so the awkward cases outnumber the easy ones. If every question has a clean answer in the corpus, the set measures nothing.
Signal
Correct source retrieved
What it shows
Whether indexing and retrieval suit the question.
Likely owner
Data or search engineering.
Signal
Citation points at the text.
What it shows
Whether a user can check the claim.
Likely owner
Product and content owner.
Signal
Answer stays in evidence.
What it shows
Whether the response controls work.
Likely owner
AI engineering
Signal
Access boundary holds
What it shows
Whether permissions survive retrieval.
Likely owner
Security and platform owner.
Signal
Safe fallback fires
What it shows
Whether users are covered when evidence is thin.
Likely owner
Product owner and operations.
Evaluate RAG failures by the control that should catch them.

Keep test examples versioned and reviewable. Feedback from production is useful. It must not become ground truth on its own. A correction may reveal a missing source, an access problem, or a reader who misread. Review it before it changes a test set, an index rule or an instruction.

A curated set of test questions checks retrieved evidence, citations and safe fallback behavior before release; reviewed user feedback and incidents become candidate test cases in a controlled loop.
Illustrative evaluation loop. Feedback becomes a reviewed improvement input rather than an automatic change to the assistant.

Security risks belong in the design

Retrieved content is untrusted input, even from an internal system. A tampered document can carry instructions aimed at the model. That is prompt injection. OWASP lists it first in its 2025 top ten for generative AI and says plainly that retrieval and fine-tuning do not fully remove it. Read the OWASP prompt injection guidance beside your own threat model.

Practical controls. Limit sources to approved collections. Hold retrieved text apart from system instructions. Restrict what tools the model may call. Gate costly actions on approval. Log enough to investigate. These cut risk. They do not make any content safe. Test hostile content and blocked-access attempts in the review set.

Buyer risk
Sensitive source leak
Question to resolve
Does retrieval enforce the source permission model?
Design response
Carry access fields into the index. Test with real roles.
Buyer risk
Stale policy answer
Question to resolve
How does a withdrawn source leave the index?
Design response
Define deletion, the source owner, and a visible freshness state.
Buyer risk
Unsupported answer
Question to resolve
What happens when evidence is missing?
Design response
Cite only when evidence is enough. Else clarify, defer or route.
Buyer risk
Prompt injection via content
Question to resolve
Can retrieved text trigger tools?
Design response
Treat content as data. Constrain tools. Test indirect injection.
Buyer risks to resolve before committing to a RAG platform.

What we saw on a real project

We run a retrieval system of our own. It answers repair questions about one vehicle model from factory manuals, forum threads and video transcripts. Every answer must cite a page or a link, or refuse. The rule in the code is one line: if no retrieved chunk supports the answer, reply “not covered”.

What was measured
Records in the index
Figure
6,264
Period or check
11 manuals, 33 web pages, 13 video transcripts, checked 22 Sep 2026
What was measured
Test questions answered with a citation
Figure
20 of 20
Period or check
golden set, run 22 Sep 2026, target was 18
What was measured
Slang and market-name aliases that expand a query
Figure
23
Period or check
added as real queries missed, 25 Aug to 1 Sep 2026
What was measured
Queries logged that returned nothing
Figure
0 of 178
Period or check
26 Aug to 22 Sep 2026
Our own retrieval system, counted from its database and test runner on 22 September 2026. Source: LC80 Brain, an Autonomous product.

Two decisions carried it. The first was to skip an embedding model. The manuals are keyword-friendly, so plain full-text search with a ranking score passed the test set. Embeddings wait until the query log shows misses that keywords cannot catch. The second was to write the 20 test questions before the search code. The runner prints a pass or fail per question, so a change that breaks retrieval is caught the same day.

What it does not prove: one corpus, one owner, no permission model. A company assistant needs the access boundary this one never had. We built that boundary for Memox, a multi-tenant product.

What breaks, and who owns it when it does

The pipeline is rarely what fails. The content under it is, and no retrieval design fixes that.

Three gaps do most of the damage. The first is a source with no owner. Nobody withdraws the March procedure, so it stays retrievable forever. The second is a permission model that lives only in the source system. Retrieval then runs as the app, not as the person. A customer finds the leak before a test does. The third is a test set one engineer wrote once and nobody has opened since. Quality drifts with no alarm on it.

Each gap needs a name against it. The content owner decides what is published and what is withdrawn. The index follows inside a stated window. The platform owner proves the permission model survives retrieval, with role tests rather than assurances. The product owner keeps the review set and decides what the assistant does when evidence is thin.

If you cannot fill in those three names today, that is the work.

What to buy, build or postpone

A managed service can speed up intake, indexing and operations. It does not set your source boundaries or your test standard. Build your own when you need odd connectors, your own permission model, or deployment controls. Fix the content first when the real problem is that nobody wrote it down.

Before you sign, ask for a proof of operation on real, approved documents and roles. Test source updates, document removal, role differences, vague questions, citations, and a source carrying indirect prompt injection. Record the limits, not the demo.

Your next step is one collection and thirty questions

Pick the one collection that already has a named owner. Write the thirty questions from the table above against it, including the ones the assistant must refuse. Run them before you write retrieval code, as we did with our 20. Keep the answers where the next person will find them. That set, not the platform, tells you when the assistant stops being trustworthy.

RAG pipelines · Project enquiry

Make your RAG system answer reliably

Tell us what the system should retrieve, cite, or protect. We will follow up to discuss whether Autonomous can help.

We use these details to respond to this enquiry. See our privacy policy.

Common questions

Does RAG stop fabricated answers?

No. Retrieval can give a model relevant evidence, but the model can still misunderstand, omit or add unsupported content. Evaluate answer support, show citations and design a safe path for weak evidence.

How much content should we index first?

One collection that is current, owned and useful for a named task. A narrow corpus makes access, freshness and test failures visible.

Who is accountable when the assistant quotes a withdrawn document?

The content owner, for pulling it at source inside a stated window. The platform owner, for the index following. If neither name exists, it will keep quoting it.

Filed under

ai-engineeringgenerative-aiknowledge-managementrag
Continue reading
Dark poster of grey dots thinning as they pass five coloured vertical bars, with four gold dots continuing on the right, illustrating leads scored by five questions.

AI & Search

How we used Jev for AI lead scoring on 68,000 Upwork jobs

We sent 211 Upwork proposals in six months and not one carries a Hired status. This is the story of what we did about it: the five questions we wrote down, the first run that cost eight cents, the test that showed our own bids were the problem, and the card the team now sees every morning.

Rizwan QaiserSep 25, 2026
12 min read
Loading page