Jon Gaudette
Draft№ 00209 MINevalsproduction

The eval you skipped is the outage you’re about to have

Every team I talk to has the same gap. They shipped the model. They never built the thing that tells them it broke.

Every team I talk to has the same gap. They shipped the model. They never built the thing that tells them it broke.

It is not laziness. Evals are the least demo-able work in the entire stack: nobody claps for a test harness, and the week you spend on one is a week the roadmap notices you were gone.11The demo is the thing that gets funded. The harness is the thing that keeps the demo true six months later. So the harness gets deferred, and the first real signal that quality moved arrives from a customer, in a support ticket, phrased as an accusation.

Start smaller than you think

The version that works is embarrassingly small. Twenty cases you collected by hand, a fixed model version, and a script that fails loudly. Take the cases from wherever you already know it slips:

  • the inputs that made somebody file a ticket last quarter
  • the formats your users paste in that nobody designed for
  • the thing it got right in March that no one has checked since
# the smallest useful regression check
for case in golden/*.json; do
  run --model "$MODEL" --input "$case" \
    | diff -q - "expected/$(basename "$case")" \
    || echo "REGRESSION $case"
done

Run it before every prompt change and every model bump. That is the whole product. You can add scoring, judges, and dashboards later, and you will — but a diff that runs is worth more than a judge that does not.


The cost of this advice is real: it is a week you would rather spend on features, and the first month it will mostly tell you things you already suspected. Pay it anyway. The alternative is learning about regressions on the schedule your users choose rather than the one you do.

How would your team find out today if quality dropped ten percent?