AI Engineering

OpenAI and ChatGPT-based chat systems

OpenAI is usually the first provider a team reaches for, and that is a reasonable default because the surrounding tooling is in one place: chat, voice, embeddings, moderation and tuning behind a single account. The risk that comes with it is a system quietly built around one vendor with no way out.

The problem

The default choice is still a choice, and it should be made explicitly

Adopting this family is rarely a mistake. The failure is that the selection is made by default and never revisited, so the architecture accumulates around one provider's specific APIs, its specific embedding model and its specific moderation categories. Two years later the assistant is one component and the coupling is everywhere, which means a price change, a deprecation or a terms change is an event the business cannot respond to. The engineering answer is not to avoid the ecosystem — it is to use it deliberately, and to keep the seams where a substitution would actually matter.

  • An assistant designed around the API rather than around the requirement
  • Voice handled as text chat with audio attached, producing an assistant that is awkward to talk to
  • Moderation decisions embedded in application code with no stated policy behind them
  • Chat, embeddings and moderation all sourced from one provider with no exit designed
  • Prompt and model changes applied without a regression set, so quality is assumed rather than known
  • Cost rising with conversation length because nothing is cached, batched or bounded
Who this is for

The people who usually bring us this problem

CTO / VP Engineering

You want to move quickly using an ecosystem your team already knows, and you want the coupling risk managed rather than left for the next person to discover.

Head of Product

The assistant has to feel native to the product — including on voice — and the quality of that experience is a design question before it is an API question.

Head of Support / CX

You need the assistant to answer from your content and hand over cleanly, and you need the safety behaviour to be defensible when someone asks how it was decided.

What it costs

What this costs while it goes unfixed

Engineering faults are rarely confined to the engineering layer. These are the commercial consequences we see most often.

Breadth is an advantage until it becomes coupling

One vendor for chat, embeddings, moderation and voice removes a great deal of integration work. The same concentration means every one of those components moves when the vendor moves, and the cost of that only becomes visible on the day something changes.

Voice is a different product, not a different transport

A voice assistant that is a chat endpoint with audio wrapped around it produces the pauses and dead air that make people distrust it. Latency, interruption and turn-taking are the design constraints, and none of them is addressed by the text stack.

Moderation without a stated policy is an undocumented decision

If the categories and thresholds are whatever the default was, the business has a policy it has never seen, applied inconsistently, and no basis on which to defend an outcome.

What we do about it

Capabilities

Each of these is work we carry out, not an area we advise on.

Chat integration built to a requirement

Integrating the chat surface against what the assistant has to do — grounded answers, defined scope, measurable resolution — rather than starting from what the API makes convenient.

Function calling into internal systems

Letting the assistant read an order, an account state or a ticket through validated tool interfaces. The arguments a model produces are treated as untrusted input at the boundary, because that is what they are.

Embeddings and semantic search

Building retrieval over your content using the embedding surface, with an index owned and refreshed by a defined process — and with a recorded decision about what happens when the embedding model changes.

Moderation with a stated policy

Deciding what the assistant must refuse, on what basis and with what user-facing outcome, then implementing it. The policy is written down, because an undocumented threshold is a decision nobody can review.

Voice and realtime assistants

Building speech interactions against their own constraints: interruption, barge-in, latency budgets, and the fact that a spoken answer must be shorter than a written one. A text assistant ported to audio reads as a bad phone menu.

Fine-tuning where it earns its keep

Usually it does not. When the requirement is tone, output-format stability or narrow classification, a tuned smaller model can be cheaper and more consistent than a prompted larger one. When the requirement is knowledge, fine-tuning is the wrong instrument entirely.

Evaluation and observability across the stack

One regression set covering chat, tools and retrieval together, plus instrumentation over cost, latency and failure rate. An assistant assembled from several components fails at the seams, so the evaluation has to cross them.

A substitution plan, designed rather than hoped for

A thin provider boundary, prompt and model versions under control, and a documented answer to what would have to change if the provider did. Cheap at design time, expensive to retrofit.

How we work

Engineering methodology

The sequence is deliberate. The order is usually what determines whether the work holds or has to be repeated.

  1. Separate the requirement from the ecosystem's shape

    Write down what the assistant must do before looking at what the platform offers. Ecosystems are persuasive — their primitives suggest architectures — and a requirement written afterwards tends to describe whatever the platform made easiest.

  2. Decide the moderation posture deliberately

    What must never be produced, what must never be accepted, and who owns that judgement. Implemented as configuration over a stated policy, so a threshold change is a reviewable decision rather than a code edit nobody sees.

  3. Design voice separately or not at all

    If a voice channel is in scope it gets its own latency budget, its own response-length rules and its own evaluation. Treating it as an output format is what produces assistants that work in a demo and are abandoned in a car.

  4. Test the seams, not only the components

    Most production failures on these systems happen between components — a tool result the prompt mishandles, a retrieval miss the model then fabricates around, a moderation outcome the application does not handle. Test the combinations, because that is where the defects are.

  5. Keep the exit cheap from the first commit

    Provider calls behind one interface, prompts in version control with the evaluation results that justified them, and no vendor-specific identifier escaping into the application's own data model.

Deliverables

What an engagement produces

Documentation is a deliverable, not an afterthought. On most of these engagements a large part of the value is a defect report precise enough for another team to act on.

Design

  • Requirement specification written before platform selection
  • Moderation and refusal policy, written down and owned
  • Voice latency and conversation design, where in scope
  • Provider boundary and substitution plan

Implementation

  • Chat integration with grounded answers and citations
  • Tool interfaces with validated arguments
  • Retrieval index with an owner and a refresh process
  • Moderation implemented as policy over configuration
  • Voice channel built to its own latency budget

Operation

  • Regression set spanning chat, tools and retrieval
  • Cost and latency instrumentation per conversation
  • Prompt and model versioning with recorded evaluation results
  • Review of real conversations for grounding and moderation failures
Under the hood

Architecture and technology

System concerns

  • Provider boundary isolating vendor-specific calls
  • Retrieval over owned content, with freshness handled
  • Tool definitions with argument validation
  • Moderation checks at input and output
  • Voice pipeline with its own latency budget
  • Conversation state and channel adaptation

Operational concerns

  • Prompt and model version control
  • Embedding model change procedure, including a re-index path
  • Cost per conversation and per channel
  • Data flow to the provider, and what is excluded
Adjacent problems

If this is not quite your problem

These overlap at the edges. Sending you to the right page is more useful than having you work it out.

Your problem is reasoning across long, complicated documents

If the answers depend on holding a policy, contract or transcript in view rather than on the surrounding tooling, choose on that basis instead.

Claude chat systems

Your problem is what one conversation costs at volume

If the assistant works and the economics do not, the provider question is about cost structure and whether the model can run on your own infrastructure.

DeepSeek chat systems

Your problem is that answers are not grounded in your content

If the assistant is plausible rather than correct, the defect is almost always in retrieval rather than in the model.

RAG and open-weight chat systems
Questions

Frequently asked

Is OpenAI the right choice for our assistant?

Often yes, and for a specific reason: if the requirement spans text, voice, embeddings and moderation, one vendor removes a large amount of integration work. The question worth asking is whether all of those are genuinely in scope. When only chat is needed, the breadth is paid for without being used, and a narrower choice can be cheaper and simpler.

Should we fine-tune on our own data?

For knowledge, no — fine-tuning teaches form rather than facts, and a model tuned on your documentation will still be wrong about it while sounding more confident. It earns its place for tone, output-format stability and narrow classification. Where the problem is that the assistant does not know your content, the fix is retrieval, not training.

Do you have an OpenAI partnership?

No. We are not a partner, reseller or certified agency for any model provider, and we take nothing from the selection. If a different provider fits better we will recommend it, which is easier to do when there is nothing to lose by saying so.

Can you build a voice assistant?

Yes, and it is a genuinely different build rather than a channel to bolt on. Spoken conversation has a latency budget measured in a few hundred milliseconds, has to handle being interrupted, and needs answers short enough to listen to. The assistant that reads well in a chat window is frequently unusable on a phone call, so the design starts from the interaction rather than from the text.

How do you keep moderation defensible?

By making it a written policy with a named owner rather than a set of defaults nobody has read. What must never be produced, what must never be accepted, and what the user sees in each case are decided and recorded, then implemented as configuration — so a threshold change is a decision someone made on purpose and an outcome can be explained afterwards.

Bring us the problem you have not been able to fix

Describe what is happening rather than what you think the cause is. If we are not the right people for it, we will say so.