Part VII: Evaluation
"All models are wrong, but some are useful."
George E. P. Box, "All models are wrong"
Evaluation looks like it should come after the system is built. In practice it reaches backward into every training run, product gate, safety rule, and model choice. A single number redirects what teams optimize. A judge, once chosen, decides which behavior survives, and an agent benchmark reshapes the harness as much as the model. Evaluation is therefore part of the system, not a label stuck on after the fact.
Chapter 47 starts with the fragile contract behind a benchmark number: held-out data, clean harnesses, stable scoring, and the many ways those promises decay. From there, Chapter 48 asks whether the observed difference is large enough to believe. The human label gets its turn in Chapter 49, treated as a designed instrument rather than a primitive. Chapter 50 moves to model judges and preference markets, where the grader has its own biases and a position in the ranking it produces. Chapter 51 separates fluency from support by decomposing answers into claims and evidence, and Chapter 52 adds state, tools, and consequence. An agent cannot be graded only by its story about what it did; the world after the task has to be inspected. Chapter 53 ends the part by asking which measurements should block a release, trigger a rollback, or become the next regression case.
Read this part with suspicion, but not cynicism. Bad evaluation is easy to mock; useful evaluation is hard engineering. The goal is to learn which numbers deserve to move a decision, which should only open an investigation, and which are really measuring the harness that produced them.
Comments
Log in to comment