Sojourner® Guides
Test the AI before users do
AI Engineering··8 min read·Bazell
Anyone can demo an LLM feature. You wire a prompt to a text box, you show it to a stakeholder, everyone agrees it's magic, and you go to lunch. The demo is free. The demo is always free.
Shipping it means answering a question the demo never asks: what counts as wrong?
The gap between "works" and "shipped"
A traditional function has a contract. Given this input, return that output, throw on these conditions. You test it by asserting the contract.
An LLM feature has no contract by default. It has a tendency. It will be right most of the time, in a distribution you have not measured, failing in ways you have not enumerated. "It worked when I tried it" is a statement about your sample size, which was one.
The move that closes the gap isn't a better prompt. It's writing down the contract yourself — as evals.
What an eval actually is
An eval is a test whose assertion is fuzzy but whose criterion is not. The criterion is the part you have to write, and it's the part people skip because it's the hard part.
Bad criterion: "the answer should be good."
Usable criteria:
- Grounded — every factual claim traces to a retrieved source. No source, no claim.
- Schema-valid — the output parses, every time, into the type the caller expects.
- Refuses correctly — out-of-scope questions get a clean decline, not a confident guess.
- Injection-resistant — instructions appearing inside retrieved data are treated as data.
- Fast enough — p95 first token under the threshold you promised.
Each of those is checkable. Some by code, some by a judge model, some by a human on a sample. All of them by someone, on a schedule.
The uncomfortable part
Once you write these down, your feature's score stops being a vibe and starts being a number. And the first number is usually bad.
That's the point. A 94% pass rate is a fact you can defend, improve, and put in a contract. "It's pretty good" is a fact about nothing, and it will not survive contact with a customer who found the 6%.
// the assertion isn't "equals" — it's "satisfies"
const verdict = await judge({
criterion: "Every factual claim is supported by a retrieved chunk.",
answer,
sources,
});
expect(verdict.pass).toBe(true);
Regression is the real prize
Here's what evals buy you that nothing else does: the ability to change the prompt without fear.
Without evals, every prompt edit is a gamble against a distribution you can't see. You improve one case and silently break four. You will not find out until a user does.
With evals, a prompt change is a diff with a score attached. You know before you merge. That is the entire discipline, and it is the same discipline as every other kind of engineering — it just took us a minute to notice, because the demo was so good.
If you can't say what wrong looks like, you haven't finished designing the feature. You've finished designing the demo.