AI Engineering

Claude-based chat systems for document-heavy work

A long context window is an advantage only when the problem is genuinely one of holding a complicated document in view. Where that is the case, a Claude-based assistant is a strong fit. Where it is not, the same choice produces a slow and expensive system. Establishing which one you have is most of the work.

The problem

Long context is a capability, not a strategy

The most common mistake on these projects is made before any code is written: assuming that a model with a large context window removes the need to think about how the corpus is organised. It does not. A corpus of any real size does not fit in any window, it changes without warning, and carrying tens of thousands of tokens on every turn is both slow and expensive. What long context genuinely buys is the ability to reason across a long and complicated document — a policy, a contract, a transcript — while holding the rest of the conversation in view. That is a real and specific advantage, and it is worth choosing deliberately.

  • An assistant that quotes the right document and the wrong clause
  • A demonstration that reads one policy convincingly and cannot be given a second
  • Latency and cost that rise with conversation length, with nothing bounding either
  • Instruction-following that visibly degrades once more than a couple of documents are in the prompt
  • No distinction between what the assistant retrieved and what it inferred
  • Refusals arriving at the wrong moment — declining a legitimate question while answering an out-of-scope one
Who this is for

The people who usually bring us this problem

Head of Support / CX

Your team answers from a large body of policy and procedure, the answers are consequential, and an assistant that is approximately right would create work rather than remove it.

CTO / VP Engineering

Model selection has been proposed on the strength of a demo, and you need it justified against the shape of the corpus, the latency budget and the cost of running it at real volume.

Head of Product

The assistant sits inside a product and has to be accurate about rules — entitlements, eligibility, process — which makes the document corpus the central design question rather than the model.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

An almost-right document assistant is worse than a search box

A search box that returns the policy leaves the reader in charge of interpreting it. An assistant that paraphrases the wrong clause is authoritative and wrong, and the cost of that asymmetry lands on whoever relied on it.

Context is a budget, and it is spent on every turn

A design that treats the context window as free produces a system whose cost and latency grow with every message in the conversation. By the time that is visible in production it has become an architecture problem rather than a tuning one.

The provider is a procurement decision as well as a technical one

Where inference runs, under whose terms and with what retention is a question a procurement or compliance function will eventually ask. Answering it during design costs a conversation; answering it after launch costs a rebuild.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Long-context reasoning over real documents

Keeping a full policy, contract or procedure in view while the conversation develops, so the answer reflects the document as a whole rather than the paragraph nearest the question. This is the capability other choices trade away, and the reason to choose this one.

Corpus preparation and structure

Deciding what the assistant may answer from and how that material is organised, so the relevant part is reliably found. Documents with internal numbering, superseded versions and exceptions need structure imposed on them before a model can be relied on over them.

Prompt caching and cost shaping

Treating the standing portion of the prompt as a cacheable asset rather than a cost paid on every turn. Caching, batching and request shaping are what turn a workable demonstration into a system with defensible unit economics.

Citation and traceable answers

Every answer carries the source it came from, so a reader can check it and support can audit it. Citation is also what makes the no-result case visible instead of silently filled.

Tool use against internal systems

Letting the assistant call a lookup — an order, an entitlement, a ticket — rather than reasoning about data it was never given. Tool calls are validated like any other boundary, because a model choosing which endpoint to call is a permission decision.

Refusal behaviour calibrated on real questions

A considered safety posture still needs its edges set for your domain. Legitimate questions your users ask must not be declined, and questions outside the assistant's coverage must not be answered. Both halves are tested against a real question set.

Evaluation on your documents

A regression set built from real questions against the actual corpus, so a prompt or model change can be assessed rather than hoped about. On document work this is the only instrument that distinguishes an improvement from a reshuffle.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Decide whether the problem is retrieval or context

    These are different problems with different answers. If questions are answered by finding one passage in a large estate, retrieval is the right architecture whatever the model. If they are answered by reasoning across a long document, context is. Confusing the two produces systems that are slow, expensive and still wrong.

  2. Prepare the corpus before writing the prompt

    Version, authority and precedence settle first: which document wins when two disagree, which is superseded, and what the assistant should do when the corpus does not cover the question. A prompt cannot resolve an ambiguity still present in the source material.

  3. Calibrate refusal against real questions, not hypotheticals

    Both failure directions are tested — legitimate questions the assistant must answer, and out-of-scope ones it must route. Harmless-sounding invented questions prove very little about either edge.

  4. Measure cost per resolved conversation

    Per-token cost is an implementation detail. The number that decides whether the system is viable is what one genuinely resolved contact costs, including retries and escalation, and that figure is measured during the build rather than discovered afterwards.

  5. Keep the model boundary substitutable

    The integration is written so the provider sits behind a boundary rather than throughout the application. Model capability, availability and price all move, and a system that only one provider can serve has handed that movement to someone else.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Design

  • Corpus assessment: scope, authority, versioning, precedence
  • Context-versus-retrieval decision, with the reasoning recorded
  • Citation and no-result behaviour specification
  • Cost model at projected conversation volume

Implementation

  • Integration behind a substitutable provider boundary
  • Corpus ingestion with version and authority handling
  • Tool interfaces with validated arguments
  • Prompt caching and request shaping
  • Refusal and routing behaviour

Operation

  • Evaluation set maintained against real questions
  • Cost and latency monitored per conversation, not per token
  • Corpus refresh process with a named owner
  • Conversation review for grounding failures
Under the hood

Architecture and technology

System concerns

  • Corpus ingestion, versioning and authority rules
  • Context assembly for the standing instruction set
  • Retrieval where the corpus exceeds the window
  • Tool calls with validated arguments
  • Citation resolution back to source
  • No-result detection and routing

Operational concerns

  • Cached standing prompt and its invalidation
  • Cost and latency per resolved conversation
  • Data flow to the provider, and what is excluded from it
  • Evaluation over a retained question set
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

Your constraint is the surrounding tooling rather than the corpus

If the deciding factor is voice, embeddings and moderation from a single vendor rather than reasoning across long documents, the OpenAI page is the closer fit — it is built around ecosystem breadth and the coupling that comes with it.

OpenAI and ChatGPT chat systems

Your constraint is cost per conversation at high volume

If the assistant is essentially right and simply cannot be afforded at the volume it was built for, the problem is unit economics rather than document handling.

DeepSeek chat systems

Your problem is finding the answer in the first place

If questions are answered by locating one passage across a large estate, retrieval architecture decides the outcome far more than the model does.

RAG and open-weight chat systems
Questions

Frequently asked

Why Claude rather than another provider?

Because the work is document-heavy and the answers have to hold across a long, complicated source, which is where this family is strongest. That is a statement about fit rather than a ranking, and it comes with a condition: if the corpus is small or the questions are answered by finding a single passage, a long-context model is the more expensive way to solve a problem retrieval solves better. We will say so, and the choice stays reversible either way.

Can we put all our documents in the prompt and skip retrieval?

Beyond a small corpus, no — and it is worth being concrete about why. The estate usually exceeds any window, it changes, and carrying it on every turn makes latency and cost grow with the conversation rather than staying flat. What does work is putting the document that matters into context and retrieving the rest, which uses the long window for reasoning rather than as bulk storage.

Are you an Anthropic partner or reseller?

No. We hold no partnership, reseller status or certification with any model provider, and we earn nothing from the choice. A provider is a dependency selected against a requirement, and the requirement is the part we are accountable for.

How do you stop it inventing policy?

By requiring a source for every answer and treating the absence of one as a first-class outcome: when nothing in the corpus supports an answer, the assistant says so and routes to a person. That case is designed and tested rather than left to chance, because it is the case that decides whether the assistant can be trusted at all.

What happens when a cheaper model becomes good enough?

The integration sits behind a boundary, so the change is a configuration edit and an evaluation run rather than a rewrite. The evaluation set is what makes that safe — without it, moving provider is a guess, which is why we build one even when nobody has asked for a migration yet.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.