Claude-based chat systems for document-heavy work
A long context window is an advantage only when the problem is genuinely one of holding a complicated document in view. Where that is the case, a Claude-based assistant is a strong fit. Where it is not, the same choice produces a slow and expensive system. Establishing which one you have is most of the work.
Long context is a capability, not a strategy
The most common mistake on these projects is made before any code is written: assuming that a model with a large context window removes the need to think about how the corpus is organised. It does not. A corpus of any real size does not fit in any window, it changes without warning, and carrying tens of thousands of tokens on every turn is both slow and expensive. What long context genuinely buys is the ability to reason across a long and complicated document — a policy, a contract, a transcript — while holding the rest of the conversation in view. That is a real and specific advantage, and it is worth choosing deliberately.
- An assistant that quotes the right document and the wrong clause
- A demonstration that reads one policy convincingly and cannot be given a second
- Latency and cost that rise with conversation length, with nothing bounding either
- Instruction-following that visibly degrades once more than a couple of documents are in the prompt
- No distinction between what the assistant retrieved and what it inferred
- Refusals arriving at the wrong moment — declining a legitimate question while answering an out-of-scope one
The people who usually bring us this problem
Head of Support / CX
Your team answers from a large body of policy and procedure, the answers are consequential, and an assistant that is approximately right would create work rather than remove it.
CTO / VP Engineering
Model selection has been proposed on the strength of a demo, and you need it justified against the shape of the corpus, the latency budget and the cost of running it at real volume.
Head of Product
The assistant sits inside a product and has to be accurate about rules — entitlements, eligibility, process — which makes the document corpus the central design question rather than the model.
What this costs while it goes unfixed
Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.
An almost-right document assistant is worse than a search box
A search box that returns the policy leaves the reader in charge of interpreting it. An assistant that paraphrases the wrong clause is authoritative and wrong, and the cost of that asymmetry lands on whoever relied on it.
Context is a budget, and it is spent on every turn
A design that treats the context window as free produces a system whose cost and latency grow with every message in the conversation. By the time that is visible in production it has become an architecture problem rather than a tuning one.
The provider is a procurement decision as well as a technical one
Where inference runs, under whose terms and with what retention is a question a procurement or compliance function will eventually ask. Answering it during design costs a conversation; answering it after launch costs a rebuild.
Capabilities
Each of these is work we carry out, not an area we advise on.
Long-context reasoning over real documents
Keeping a full policy, contract or procedure in view while the conversation develops, so the answer reflects the document as a whole rather than the paragraph nearest the question. This is the capability other choices trade away, and the reason to choose this one.
Corpus preparation and structure
Deciding what the assistant may answer from and how that material is organised, so the relevant part is reliably found. Documents with internal numbering, superseded versions and exceptions need structure imposed on them before a model can be relied on over them.
Prompt caching and cost shaping
Treating the standing portion of the prompt as a cacheable asset rather than a cost paid on every turn. Caching, batching and request shaping are what turn a workable demonstration into a system with defensible unit economics.
Citation and traceable answers
Every answer carries the source it came from, so a reader can check it and support can audit it. Citation is also what makes the no-result case visible instead of silently filled.
Tool use against internal systems
Letting the assistant call a lookup — an order, an entitlement, a ticket — rather than reasoning about data it was never given. Tool calls are validated like any other boundary, because a model choosing which endpoint to call is a permission decision.
Refusal behaviour calibrated on real questions
A considered safety posture still needs its edges set for your domain. Legitimate questions your users ask must not be declined, and questions outside the assistant's coverage must not be answered. Both halves are tested against a real question set.
Evaluation on your documents
A regression set built from real questions against the actual corpus, so a prompt or model change can be assessed rather than hoped about. On document work this is the only instrument that distinguishes an improvement from a reshuffle.
Engineering methodology
The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.
Decide whether the problem is retrieval or context
These are different problems with different answers. If questions are answered by finding one passage in a large estate, retrieval is the right architecture whatever the model. If they are answered by reasoning across a long document, context is. Confusing the two produces systems that are slow, expensive and still wrong.
Prepare the corpus before writing the prompt
Version, authority and precedence settle first: which document wins when two disagree, which is superseded, and what the assistant should do when the corpus does not cover the question. A prompt cannot resolve an ambiguity still present in the source material.
Calibrate refusal against real questions, not hypotheticals
Both failure directions are tested — legitimate questions the assistant must answer, and out-of-scope ones it must route. Harmless-sounding invented questions prove very little about either edge.
Measure cost per resolved conversation
Per-token cost is an implementation detail. The number that decides whether the system is viable is what one genuinely resolved contact costs, including retries and escalation, and that figure is measured during the build rather than discovered afterwards.
Keep the model boundary substitutable
The integration is written so the provider sits behind a boundary rather than throughout the application. Model capability, availability and price all move, and a system that only one provider can serve has handed that movement to someone else.
What an engagement produces
Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.
Design
- Corpus assessment: scope, authority, versioning, precedence
- Context-versus-retrieval decision, with the reasoning recorded
- Citation and no-result behaviour specification
- Cost model at projected conversation volume
Implementation
- Integration behind a substitutable provider boundary
- Corpus ingestion with version and authority handling
- Tool interfaces with validated arguments
- Prompt caching and request shaping
- Refusal and routing behaviour
Operation
- Evaluation set maintained against real questions
- Cost and latency monitored per conversation, not per token
- Corpus refresh process with a named owner
- Conversation review for grounding failures
Architecture and technology
System concerns
- Corpus ingestion, versioning and authority rules
- Context assembly for the standing instruction set
- Retrieval where the corpus exceeds the window
- Tool calls with validated arguments
- Citation resolution back to source
- No-result detection and routing
Operational concerns
- Cached standing prompt and its invalidation
- Cost and latency per resolved conversation
- Data flow to the provider, and what is excluded from it
- Evaluation over a retained question set
If this is not quite your problem
These overlap at the edges. Sending you to the right page is more useful than having you work it out.
Your constraint is the surrounding tooling rather than the corpus
If the deciding factor is voice, embeddings and moderation from a single vendor rather than reasoning across long documents, the OpenAI page is the closer fit — it is built around ecosystem breadth and the coupling that comes with it.
OpenAI and ChatGPT chat systemsYour constraint is cost per conversation at high volume
If the assistant is essentially right and simply cannot be afforded at the volume it was built for, the problem is unit economics rather than document handling.
DeepSeek chat systemsYour problem is finding the answer in the first place
If questions are answered by locating one passage across a large estate, retrieval architecture decides the outcome far more than the model does.
RAG and open-weight chat systemsFrequently asked
Why Claude rather than another provider?
Because the work is document-heavy and the answers have to hold across a long, complicated source, which is where this family is strongest. That is a statement about fit rather than a ranking, and it comes with a condition: if the corpus is small or the questions are answered by finding a single passage, a long-context model is the more expensive way to solve a problem retrieval solves better. We will say so, and the choice stays reversible either way.
Can we put all our documents in the prompt and skip retrieval?
Beyond a small corpus, no — and it is worth being concrete about why. The estate usually exceeds any window, it changes, and carrying it on every turn makes latency and cost grow with the conversation rather than staying flat. What does work is putting the document that matters into context and retrieving the rest, which uses the long window for reasoning rather than as bulk storage.
Are you an Anthropic partner or reseller?
No. We hold no partnership, reseller status or certification with any model provider, and we earn nothing from the choice. A provider is a dependency selected against a requirement, and the requirement is the part we are accountable for.
How do you stop it inventing policy?
By requiring a source for every answer and treating the absence of one as a first-class outcome: when nothing in the corpus supports an answer, the assistant says so and routes to a person. That case is designed and tested rather than left to chance, because it is the case that decides whether the assistant can be trusted at all.
What happens when a cheaper model becomes good enough?
The integration sits behind a boundary, so the change is a configuration edit and an evaluation run rather than a rewrite. The evaluation set is what makes that safe — without it, moving provider is a guess, which is why we build one even when nobody has asked for a migration yet.
Related capabilities and work
Bring us the problem you have not been able to fix
Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.