Summary
Evaluation begins with a decision, not a score. A defensible evaluation states the claim, the complete system being compared, the target population, the failure costs, and what evidence would change the decision. Every score is the output of a measurement instrument, not a property of model weights. The model, harness, task manifest, scorer, and aggregation rule jointly determine the result. Pin those components, retain the case-level records, and report only the qualified claim that this evidence supports.
Once the measured object is fixed, the analysis determines how strongly the observations support the decision. It names the estimand and independent sampling unit, preserves paired cases, accounts for run and scorer variation, and asks whether the difference reaches a decision-relevant effect. A narrow interval can establish precision under the declared design, but uncertainty does not repair bias in the task sample, answer key, rubric, or scorer. More rows cannot rescue a measurement instrument aimed at the wrong construct.
A human label is an observation made under a protocol: a particular rater saw particular evidence through a particular interface and applied a particular rubric. A model judge is a versioned instrument that must be validated against suitable human judgments or checkable outcomes before it can scale them. Specialized evaluations must keep their targets distinct. Truth, grounding, attribution, and citation quality answer different questions about claims and sources. Agent evaluation goes beyond the final message to verify the state the system changed, while using trajectories to enforce only declared process constraints or to diagnose failures.
Evidence becomes operational only when its lifecycle and authority are explicit. Development, locked confirmation, and diagnostic suites serve different purposes and must not silently exchange cases. Offline, shadow, canary, and production evidence answer different questions, so one stage cannot stand in for all the others. A release gate can then promote, hold, narrow, roll back, or escalate a named system revision under predeclared margins, guardrails, ownership, and rollback rules. The evidence expires when its system, population, or measurement assumptions no longer hold. Part VIII starts from that dependency on evidence and asks what must be constrained, who has authority, and where the control is enforced.
Comments
Log in to comment