AI Engineering

DeepSeek-based chat systems where cost per conversation decides the design

There is a class of assistant where the technology works and the economics do not: the volume is high, each conversation is short, and the price per token is the difference between a system that pays for itself and one that does not. That problem is solved with cost engineering and deployment choices, and this family of models is unusually relevant to it.

The problem

A working assistant with unworkable unit economics

Assistants fail commercially far more often than they fail technically. A design that resolves contacts well at demonstration volume can be unaffordable at the volume it was built for, and no amount of prompt engineering changes that — the cost is per conversation and the conversations are the load. The lever is not a cleverer prompt; it is which model runs, where it runs, and how much of the traffic needs a large model at all. The second half of the problem is rarely raised early enough: where inference happens, under whose jurisdiction and on whose terms is a procurement question with real answers, and it constrains the deployment as much as the budget does.

  • An assistant that works and costs more per conversation than the contact it replaces
  • A procurement or compliance question about where inference runs that has never had a straight answer
  • No plan for what happens if a provider changes its pricing, terms or availability
  • Quality that is adequate and cost that is not, with the trade never examined the other way round
  • Self-hosting discussed as a principle rather than as a costed option with an operating commitment attached
  • Open weights assumed to be a drop-in replacement for a hosted frontier model
Who this is for

The people who usually bring us this problem

CFO / Commercial lead

The assistant has been approved on a business case that depends on a cost per contact, and you need that figure to be an engineering input rather than an afterthought.

CTO / VP Engineering

You need an honest account of what running a capable open-weight model on your own infrastructure involves, and where it is genuinely cheaper than paying per token.

Head of Support / CX

Your volume is high and each contact is short — the shape where routing most conversations to a smaller, cheaper model is both viable and decisive.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Unit economics decides whether the project was worth doing

An assistant that resolves contacts at a higher cost than handling them is a net loss that scales with success, which is the worst kind. Cost per resolved conversation is therefore a requirement of the design, not a metric discovered at the end.

Jurisdiction is a procurement answer, not a technical one

Where inference happens determines which laws and terms apply to the data in the prompt. Someone with the authority to block the launch will ask, and answering late means a rebuild rather than a configuration change.

Concentration is a business risk when a provider can end it

A dependency whose pricing, terms and availability are set elsewhere is one whose continuity is not yours. The mitigation is not necessarily self-hosting — but the option has to exist before it is needed.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Cost per resolved conversation

Modelling and then instrumenting the figure that decides viability: not cost per token, but what one genuinely resolved contact costs including retries, tools and escalation. The measurement the whole architecture is chosen against.

Model routing and tiering

Most contacts in a high-volume support queue are simple and a small fraction are not. Routing the simple majority to a smaller, cheaper model and escalating the rest is usually the largest lever available, and it is an engineering decision rather than a purchasing one.

Open-weight deployment on your own infrastructure

Running published model weights in your own environment, on your own hardware or in your own cloud account, with the serving stack, capacity and monitoring that requires. This is what makes cost predictable and the data-flow question answerable.

Serving capacity and the fixed-cost trade

Self-hosting converts a variable per-token bill into a fixed capacity cost. That trade only works above a utilisation threshold, needs headroom for peaks, and carries an operating commitment for as long as it runs — all established before recommending it rather than after.

Jurisdiction and data-flow design

Establishing what leaves your environment, to where, under which terms and with what retention, then designing to constrain it. Where the answer has to be that nothing leaves, that is a deployment constraint decided at the start rather than a preference expressed at the end.

Comparison on your own tasks

Open weights and hosted frontier models differ in ways general benchmarks do not capture for your workload. The comparison is run over your real tasks, scored on your definition of acceptable, because that is the only comparison that predicts anything about production.

The exit path, costed in advance

A documented answer to what moving between a hosted provider and your own infrastructure would cost, take and risk. Producing that document while nothing is wrong is the difference between a migration being an option and being an incident.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Establish the volume curve before the architecture

    Contacts per month, tokens per conversation, and how each is expected to grow. A design chosen before those numbers exist is chosen against a guess, and the two viable architectures here — hosted per-token and self-hosted fixed-capacity — diverge sharply depending on where the curve lands.

  2. Treat self-hosting as a capacity decision

    The honest comparison is cost per conversation at a given utilisation, including the operational cost of running it and the hardware headroom a peak needs. Below a threshold, paying per token is genuinely cheaper, and saying so is more useful than treating self-hosting as a position.

  3. Answer the data-flow question with a diagram

    Which fields reach the model, which are excluded before the request is built, where inference runs and what is retained. A diagram is what a procurement or compliance review can act on, and drawing it early usually simplifies the design.

  4. Compare quality on your tasks, not on published scores

    Benchmarks measure what benchmarks measure. The useful comparison runs both options over a retained set of your own conversations and scores the outputs against the standard your team already applies, including the cases where both are judged unacceptable.

  5. Keep the routing layer yours

    Whatever runs the model per request sits behind an interface we own, so mixing a cheap model for most traffic with a stronger one for the rest is a configuration rather than a rewrite — and so is changing which model holds each position.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Design

  • Cost model per resolved conversation, at projected volume
  • Hosted-versus-self-hosted comparison, with its utilisation threshold
  • Routing rules: which traffic reaches which model, and why
  • Data-flow diagram covering jurisdiction, terms and retention

Implementation

  • Routing layer behind an owned interface
  • Open-weight serving deployment where that is the right answer
  • Evaluation harness scoring both options on your own tasks
  • Cost and latency instrumentation per conversation and per route

Operation

  • Capacity monitoring with peak headroom tracked
  • Routine cost review against the original model
  • Model update procedure, including the re-evaluation it requires
  • Documented exit path between hosted and self-hosted
Under the hood

Architecture and technology

System concerns

  • Request classification and routing
  • Model endpoints: hosted, self-hosted, or both
  • Prompt construction per route, with caching
  • Tool interfaces with validated arguments
  • Escalation from the cheap route to the strong one
  • Fallback when a route is unavailable

Operational concerns

  • Capacity, utilisation and peak headroom on owned hardware
  • Cost per conversation tracked against the model's assumption
  • Data flow, jurisdiction and retention, reviewed
  • Update procedure for weights and serving stack
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

Your problem is document handling rather than cost

If the answers depend on reasoning across long source documents and the budget is not the binding constraint, choose on that basis instead.

Claude chat systems

Your problem is the ecosystem around the model

If having voice, embeddings and moderation in one place is what matters, that is a different selection problem.

OpenAI and ChatGPT chat systems

Your problem is running the model yourself, whatever it is

Self-hosting an open-weight model is only half of a deployment. Grounding it in your own corpus is the other half, and it is usually where the quality actually comes from.

RAG and open-weight chat systems
Questions

Frequently asked

Is a cheaper model actually good enough?

For a large share of a typical support queue, frequently yes — short, well-scoped questions with a clear answer in your own documentation. For the rest, no, and pretending otherwise produces confident wrong answers at the exact moment a customer is most invested. That is precisely why the routing layer matters: it lets the answer be "both, chosen per request" instead of committing the whole system to one side of a trade.

Does self-hosting mean our data never leaves?

It means the inference happens in your environment, which is the material part — but it is not the whole data-flow picture. Retrieval indexes, logging, evaluation and support access all touch the same data, and a competent answer describes every path rather than only the model call. We map all of it, because the compliance review that follows will ask about all of it.

Will open weights match a hosted frontier model?

On hard reasoning, usually not, and anyone claiming otherwise has a comparison to sell. The relevant question is narrower: for your traffic, on your tasks, does the difference fall inside or outside what your customers can tolerate? That is answerable, by scoring both options on a retained set of your own conversations, and it is a much better basis for the decision than a benchmark table.

Do you have a relationship with DeepSeek?

No, and the same is true of every model provider we work with. We hold no partnership or reseller arrangement and we take nothing from the selection, which is what makes it possible to recommend against a provider when the requirement points elsewhere.

What does running our own inference actually require?

Hardware or cloud capacity sized for peak rather than average, a serving stack maintained through its own updates, monitoring, and someone accountable for it on an ongoing basis. It is a real operating commitment rather than a one-off build, which is why it is recommended against a utilisation figure instead of as a general principle. Below that threshold, paying per token is the cheaper and simpler answer and we will say so.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.