Building reliable AI products beyond the prototype
A demo that works on ten hand-picked inputs tells you almost nothing about the ten-thousandth real one. Here is how we close that gap.
4 min read
Every AI project we take on has a moment, usually in the first two weeks, where the prototype works. Someone pastes a real document into a notebook, the model returns exactly the right answer, and the room gets excited. That moment is worth enjoying. It is also the least informative result the project will ever produce.
A prototype is tested on inputs its builders chose. Production is tested on inputs nobody chose: blurred scans, half-French half-English emails, a PDF that is secretly an image, a customer who writes in capital letters. The work of shipping AI is the work of closing the gap between those two distributions.
The prototype lies politely
Large language models fail differently from traditional software. A broken parser throws an exception. A model under pressure returns something fluent, well-formatted and wrong. Nothing crashes, nothing alerts, and the error travels downstream looking exactly like a correct answer.
That is why we treat "it worked when I tried it" as a hypothesis rather than a result. The question is never can the model do this task. The question is how often, on which inputs, and what happens when it doesn't.
Start with an evaluation set, not a prompt
Before we tune a single prompt, we build the set of examples we will judge every prompt against. It does not need to be large. Two hundred well-chosen cases beat two thousand random ones. What matters is that it looks like production:
- Real inputs, anonymised, pulled from the client's actual workflow — not examples written by the team.
- Both languages in the proportions users actually send them. In Cameroon that usually means French, English, and a healthy amount of both in one message.
- Known hard cases: low-quality scans, missing fields, contradictory information.
- Cases where the right answer is "I don't know." A system that never abstains is guessing.
Then we make the evaluation cheap to run, so it runs on every change:
type Case = { id: string; input: string; expected: ClaimFields | "needs_review" };
export async function runEval(cases: Case[], extract: Extractor) {
const results = await Promise.all(
cases.map(async (c) => {
const output = await extract(c.input);
return { id: c.id, pass: matches(output, c.expected), output };
}),
);
const passRate = results.filter((r) => r.pass).length / results.length;
return { passRate, failures: results.filter((r) => !r.pass) };
}
The pass rate matters less than the failures list. Reading twenty failures in a row teaches you more about a system than any dashboard.
Make failure a first-class output
Most AI features are designed for the happy path and then patched when reality arrives. We design the unhappy path first. Every model output is parsed against a schema, and the schema includes a way to say this needs a person:
const Extraction = z.discriminatedUnion("status", [
z.object({ status: z.literal("complete"), fields: ClaimFields, confidence: z.number() }),
z.object({ status: z.literal("needs_review"), reason: z.string() }),
]);
If the model's response does not parse, that is not an edge case to log and ignore — it is a needs_review result like any other, routed to the same queue. The downstream system never has to wonder whether it received a real answer.
A model is a dependency that can change its behaviour without a version bump. Pin what you can, measure what you can't, and never let an unparsed answer reach a customer.
Put people where the uncertainty is
"Human in the loop" often means a person approving everything, which is expensive and quickly becomes a rubber stamp. We try to put people exactly where the system is uncertain and nowhere else.
In practice that means routing on confidence and on consequence. A low-confidence extraction on a small claim goes to review. A high-confidence extraction on a very large claim also goes to review, because the cost of a rare mistake is high. Everything else flows straight through, and a random sample of it is reviewed weekly so the confidence scores stay honest.
Observe everything in production
Once real traffic arrives, the evaluation set starts to age. New document formats appear, a partner changes their template, users discover a use nobody planned for. So we log, for every call:
- The input (or a reference to it), the prompt version and the model version.
- The raw output, the parsed output, and whether parsing succeeded.
- Latency, token cost and the route the result took.
Every week, reviewed production failures become new evaluation cases. The evaluation set grows in exactly the places the system is weakest, which is the only place growth is useful.
The checklist we ship with
None of this is exotic. It is ordinary engineering discipline applied to a component that happens to be probabilistic:
- An evaluation set drawn from production, run on every change.
- Schema-validated outputs with an explicit needs review state.
- Confidence- and consequence-based routing to people.
- Versioned prompts, pinned models, and logs that tie every answer to both.
- A weekly loop that turns production failures into test cases.
The prototype tells you the idea is possible. This checklist is what makes it dependable — and dependable is what people will actually use.