The figure on the slide has usually been justifying decisions for a year. When we rebuild the evaluation set properly, it moves. Here is why that keeps happening and how to build a held-out set that actually tells you something.
There is a number on a slide. It is usually a percentage in the low nineties, it appeared during the pilot, and it has been quoted in every subsequent conversation about whether the system can be trusted with something new.
We ask three questions about it. Where did the test data come from. Who chose which examples went in. And can we see the set. In a meaningful proportion of engagements, one of those three questions does not have a clean answer.
Nobody sets out to measure a model on its own training data. It happens through ordinary decisions made under time pressure.
None of these is misconduct. Each is a small, defensible decision, and together they produce a number that is measuring something other than what everyone believes it is measuring.
We do not accept a client's evaluation set, and we do not accept a vendor's. Not because we distrust the people involved, but because the whole value of the exercise depends on the set being genuinely unseen, and neither of those parties can prove that to us or to themselves.
We build from material the system has not been exposed to, weighted towards the cases that actually matter to the business rather than the cases that are easy to generate. Then we run it, and the resulting figure goes into the report next to the one on the slide.
The single percentage is the least useful output of a proper evaluation. What the business needs is the shape of the failures.
Does it fail more on long inputs. On a particular customer segment. On the cases where being wrong is most expensive, which is depressingly common because those cases are usually rare and therefore under-represented in whatever the team assembled. Does confidence correlate with correctness at all, which determines whether a human review threshold is doing anything useful or is theatre.
A system at eighty seven percent that fails predictably and cheaply is a better system than one at ninety three that fails unpredictably on the expensive cases. The headline number cannot tell you which one you have.
An evaluation is a photograph. Inputs change, the underlying model gets updated by a provider, the retrieval index fills up with new material. A figure measured eighteen months ago is a historical fact rather than a current one.
The fix is unglamorous. Keep the held-out set, run it on a schedule, and record the result somewhere a governance committee can see it. Most organisations we assess have never run their evaluation twice.
Find out where the number came from. If nobody can tell you, that is your answer and it is worth knowing. If they can, ask whether any of that test material could plausibly have reached the model through tuning or retrieval.
Then hold back some genuinely new material, weighted towards expensive failure cases, and run it once. Even a rough version of this exercise will tell you more than the slide does, and it costs a couple of weeks rather than an engagement.
We build the evaluation set, run it, and put the real figure next to the one on your slide.