AI Infra
0%
Summary

Summary

AuthorChangkun Ou
Reading time~1 min

Evaluation looked like an external score only until the part followed where scores are used. A benchmark changes training incentives. A statistical test decides whether a difference is real enough to matter, while a rubric shapes the human label underneath it. The model judge brings its own bias; factuality work decomposes claims and evidence. Agent evaluation has to inspect the changed world, and operational evaluation, last in the chain, decides whether a release moves forward.

The main concern is that measurement is easy to corrupt without noticing. Data leaks into the training set, tasks saturate, a sample stays too small to support the claim it carries, judges can be gamed, and a benchmark drifts until it is measuring the harness that produced the result rather than the model. A number should therefore carry an authority level before it blocks a release, drives training, or changes a product decision.

The reader should leave with suspicion, but not cynicism. Useful evaluation is hard engineering, not a decorative leaderboard. The open problem is measuring long-horizon tasks, workplace productivity, and agent side effects before the instruments saturate or become targets of optimization. Part VIII starts from that dependency on evidence and asks what should be constrained, who has authority to constrain it, and where a safety or governance claim can actually be enforced.

Comments

Log in to comment