AI Engineering

RAG and open-weight chat systems, grounded in your own corpus

When an assistant is confident and wrong, the model is almost never the cause. The answer was not in what it was given, or the relevant material was not retrieved, or it was retrieved and the prompt buried it. Retrieval is where these systems are won, and it is usually built last.

The problem

The failure is in retrieval, and retrieval is built after everything else

Retrieval-augmented generation is the default architecture for assistants that answer from company content, and it is usually assembled in the wrong order: the model, the prompt and the interface come first, and retrieval is bolted underneath at the end. That ordering guarantees the hardest part gets the least attention. When a chunking strategy severs a clause from the table it belongs to, when an exact product code fails to match because only semantic similarity is searched, or when a document is returned to someone not entitled to see it, no prompt improvement recovers any of it. These are retrieval defects. They are diagnosable, cheap to prevent and expensive to discover in production.

  • Answers that are plausible, well-written and contradicted by the source they claim
  • Exact identifiers — codes, references, error numbers — that never retrieve anything
  • Documents split in a way that separates a condition from the exception that qualifies it
  • An index months out of date with no owner and no refresh process
  • No way to tell whether an answer failed because retrieval missed or because generation drifted
  • Material returned to a user who is not authorised to read it
  • Answers that depend on one vendor's availability, terms and pricing with no alternative owned
Who this is for

The people who usually bring us this problem

CTO / VP Engineering

You are building retrieval over your own content and need the parts that determine answer quality — chunking, hybrid matching, reranking, evaluation — built competently rather than assumed.

Head of Support / CX

You need answers sourced from documentation your team maintains, with a citation, so support can see what the assistant relied on when it was wrong.

CTO / Compliance lead

Content is permissioned, and an assistant that retrieves across the whole corpus regardless of who is asking is not a quality problem but a disclosure one.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Retrieval defects are invisible from the outside

A bad chunking decision produces fluent, confident, wrong answers. Nothing in the interface distinguishes those from correct ones, so the defect surfaces as customers losing trust rather than as an error anyone can see.

Permissioned content turns a retrieval bug into a disclosure

If retrieval searches everything and filtering happens afterwards — or does not happen at all — the assistant will eventually show someone a document they were never entitled to read. That is an access-control failure wearing a chat interface.

An unowned index decays quietly

Retrieval quality tracks index freshness. Without an owner and a process, the assistant's answers are correct on the day it launches and gradually less so, with no step change that prompts anyone to investigate.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Retrieval design: chunking with structure preserved

Splitting documents so a clause keeps its heading, a row keeps its table and an exception stays with its rule. Most retrieval failures on policy and technical content begin here, which is why it is examined first rather than last.

Hybrid search and reranking

Semantic similarity alone cannot match an exact part number, and keyword matching alone cannot find a paraphrase. Running both and reranking the combined result is what makes an assistant reliable across both kinds of question — which real users ask in the same session.

Permission-aware retrieval

Filtering by entitlement inside the retrieval step rather than after it, so content a user is not authorised to see is never assembled into a prompt. Designed from the source system's own access model rather than a parallel one invented for the assistant.

Embedding selection and change handling

Choosing an embedding model against the content and the question types, and deciding in advance what happens when it changes, because changing it means re-indexing everything. That is a scheduled operation, not an incident.

Index ownership and freshness

A named owner, a refresh process, and monitoring that reports staleness rather than leaving it to be discovered from a wrong answer. Retrieval quality is a property of the index's currency as much as of its construction.

Evaluation that separates retrieval from generation

Scoring whether the right passage was retrieved independently of whether the answer was good. Without that split, every failure looks like a prompt problem, which is why teams spend months tuning prompts against defects that were never there.

Open-weight models served on your own infrastructure

Running published weights in your own environment where cost predictability, data residency or continuity matter. The model is one component; the serving stack, capacity and its ongoing operation are the commitment that comes with it.

No-result behaviour, designed first

What happens when retrieval returns nothing useful — the single case that decides whether the assistant can be trusted. It says so and routes, rather than answering from the model's own priors, which is where confident fabrication enters.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Measure retrieval before touching the prompt

    For a set of real questions, establish whether the correct passage is retrieved at all and where it ranks. If it is not retrieved, prompt work is wasted effort; if it is retrieved and still not used, the problem is downstream. This single step reorders the whole project, which is why it is done first.

  2. Build the evaluation set from real questions

    Questions people actually asked, including the ones phrased badly, asked in two parts, or using internal shorthand. A set written by the team building the system tests the system's assumptions rather than its behaviour.

  3. Design the no-result path before the happy path

    It is the case that determines trustworthiness, and it is trivial to add at design time and awkward to retrofit once the interface and the conversation flow assume an answer always exists.

  4. Treat the index as a product with an owner

    A refresh process, a staleness check, and a person accountable for it. This is unglamorous and it is the difference between an assistant that is right for a quarter and one that is right for years.

  5. Make the serving stack an explicit commitment

    Where the model runs on owned infrastructure, that is an operating commitment with capacity, updates and monitoring attached rather than a one-off deployment. It is stated as such, including when the honest recommendation is to keep paying per token instead.

  6. Keep the model replaceable

    Retrieval quality is portable between models; a prompt tuned to one model's idiosyncrasies is not. Building so the model sits behind a boundary means the retrieval investment survives every subsequent model change.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Design

  • Corpus inventory with authority and permission structure
  • Chunking strategy derived from the content's own shape
  • Hybrid retrieval and reranking design
  • Evaluation set built from real user questions

Implementation

  • Index build with permission metadata carried through
  • Hybrid search with reranking
  • Permission filtering inside retrieval rather than after it
  • Citation resolution to source documents
  • No-result detection and routing
  • Open-weight serving, where that is the deployment choice

Operation

  • Index refresh process with a named owner
  • Staleness monitoring
  • Retrieval and generation scored separately, on a retained question set
  • Capacity and cost monitoring where the model is self-hosted
Under the hood

Architecture and technology

System concerns

  • Ingestion with structure and permission metadata preserved
  • Chunking that respects document hierarchy
  • Hybrid retrieval: semantic and exact matching
  • Reranking over the combined candidate set
  • Permission filtering inside the retrieval step
  • Citation resolution and no-result routing

Operational concerns

  • Index freshness, ownership and staleness reporting
  • Embedding model change and the re-index it requires
  • Retrieval and generation evaluated separately
  • Serving capacity where the model runs on owned infrastructure
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

Your questions are answered from long documents rather than a wide corpus

If the answer depends on holding a whole policy or transcript in view and the estate is small, retrieval may be the wrong architecture and a long-context choice the right one.

Claude chat systems

Your constraint is the surrounding tooling

If you need voice, moderation and embeddings from one vendor, that is a different selection problem from retrieval quality.

OpenAI and ChatGPT chat systems

Your constraint is what a conversation costs

If retrieval is working and the economics are not, the question moves to deployment and routing.

DeepSeek chat systems
Questions

Frequently asked

Do we need RAG at all?

Fewer teams need it than build it. If the corpus is small enough to place in context, if the questions are answered by one passage rather than by reasoning across material, or if a structured lookup would do the job, retrieval adds a dependency and a failure mode for nothing. It earns its place when the estate is larger than any context window, when content changes, or when an answer must cite its source — and those are the conditions we check before designing one.

Which open-source models do you use?

Whichever the requirement points at, weighed on the client's own tasks rather than on published leaderboards, and chosen alongside the deployment question rather than after it. The more material point is that retrieval quality is largely independent of this choice, which is why we build that part first and keep the model behind a boundary.

How do you stop the assistant showing someone a document they should not see?

By filtering inside retrieval rather than after it, using the source system's own permission model rather than a parallel one invented for the assistant. If content is assembled into a prompt and filtered afterwards, the exclusion is a promise about the prompt rather than a property of the pipeline — and the prompt is the wrong thing to make promises about.

Why is our assistant confidently wrong when the document is right there?

Almost always because the passage was never retrieved, or was retrieved below the material that displaced it. The way to know which is to score retrieval separately from generation: check whether the correct passage appears in what the model was given, before assessing what it did with it. Teams routinely spend weeks rewriting prompts against a retrieval defect, and this measurement is what prevents that.

Is a self-hosted model cheaper?

Above a utilisation threshold, usually; below it, usually not. Self-hosting turns a variable per-token cost into fixed capacity, which only pays off when the capacity is genuinely used, and it adds an operating commitment that outlives the project. We work it out against your projected volume and tell you which side of the line you are on, including when the answer is to keep paying per token.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.