AI Infra
0%
Summary

Summary

AuthorChangkun Ou
Reading time~1 min

Post-training began with the base model's missing interface. Demonstrations, adapters, behavior specifications, preference comparisons, reward models, direct preference objectives, verifiable rewards, safety policies, and synthetic data loops all try to say what kind of behavior should survive. They differ in optimizer and pipeline, but they share one dependency: some signal has to tell the model what better means.

That signal is the fragile part. When it is underspecified, the model learns style instead of substance. An exploitable signal turns the reward into a target to game. And if the judge is weak, the loop can overfit it and still look improved. This is why the part treats alignment less as a slogan than as governance over signals, reviewers, filters, verifiers, and policy text.

The reader should carry forward a simple test: ask who wrote the signal, how it is checked, where it fails, and when the system is allowed to trust its own output. The open question is how far synthetic data, AI feedback, and verifiers can push improvement before the judge, filter, or policy becomes the new ceiling. Part IV asks what can be left to inference time instead of being frozen into weights: search, verification, tool calls, traces, and extra computation spent only when a request seems to need it.

Comments

Log in to comment