Skip to content
All insights
AI & ML4 min

Evaluating LLM features before you launch them

Shipping an AI feature without an evaluation harness is shipping a system nobody can prove works. Here is the eval setup we build on every engagement.

Every LLM feature we take to production ships with an evaluation harness before it ships with a UI. Not because it is best practice on a slide, but because an LLM feature has no compiler, no type system and no deterministic test to tell you it still works after the next prompt tweak or model upgrade. The harness is the only thing standing between "it looked fine in the demo" and a regression your users find first.

Why unit tests are not enough

A traditional test asserts that an input maps to an exact output. LLM output is non-deterministic, phrased differently on every call, and correct in more than one way. Asserting string equality gives you a test suite that is red for reasons that do not matter and green for reasons you cannot trust.

The shift is from "is this output equal to the expected string" to "does this output satisfy the properties we care about" — is it grounded in the retrieved context, is it in the right format, does it refuse when it should, is it within a latency and cost budget. Those are graded, not asserted, and the grade is a distribution you watch over time.

Building the golden dataset

We start with 50 to 200 real examples drawn from the actual use case: support tickets, contract clauses, product questions — whatever the feature operates on. Each example gets an input, any context the system would retrieve, and a human-written note on what a good answer must contain. Synthetic data fills gaps for rare cases, but the core set is real traffic.

The dataset is versioned in the repository next to the code. When a stakeholder says "it got this wrong," that example becomes a new row. The set only grows, and every failure that reaches production is represented in it before the fix is merged.

Graders: code, model, and human

Cheap deterministic checks run first — JSON parses, required fields present, no PII echoed, response under the token cap. These catch the majority of real regressions and cost nothing.

For the subjective dimensions we use an LLM-as-judge with a rubric: a separate model call scores groundedness and helpfulness on a fixed scale with the rubric in the prompt. It is not perfect, but it is consistent, and we calibrate it against a human-labelled sample every few weeks.

A small slice — usually 20 examples — is always reviewed by a person each release. The model judge tells you the direction; the human review tells you whether to trust the judge.

Wiring it into CI and production

The harness runs on every pull request that touches a prompt, a retrieval parameter, or the model version. A drop in the aggregate score below the agreed threshold blocks the merge the same way a failing test would.

The same graders run on a sample of live traffic. Production scores are the early warning that a model provider changed something underneath you, or that real inputs have drifted away from your dataset. When the two diverge, the dataset needs new rows.

What this costs and what it buys

A first harness is two to four days of work: assemble the dataset, write the graders, wire the runner. It is the cheapest insurance you will buy on an AI project.

The return is the ability to change things. Upgrade the model, cut the prompt in half, swap the vector store — and know within minutes whether quality held. Without the harness every one of those changes is a gamble you resolve in production.

30 minutes · CET · no deck

Book the call. Leave with a plan.

Bring the paper, the WhatsApp thread, or the spreadsheet. We will tell you what to build first — and what not to.