AI Infra
0%
Summary

Summary

AuthorChangkun Ou
Reading time~1 min

The opening part did not try to teach the whole stack at once. It followed one request across the layers, then pulled back to mark which claims are settled enough to build on and which still belong to live argument. It cleaned up the vocabulary the field borrows from neighboring disciplines, because a useful metaphor can become a misleading mechanism if it is read too literally. And it drew the subject's boundary: the generative lifecycle this book covers sits beside the ranking and recommendation infrastructure that already ran the world, inheriting its retrieval funnels and measurement discipline while needing a serving tier of its own.

Local symptoms often have non-local causes. A benchmark jump, a latency incident, a large GPU bill, or an agent failure may be decided by a layer the symptom does not name. What should survive this part is a habit, not a taxonomy: when a design choice seems arbitrary, look one layer down and ask who is paying the cost.

The open work is keeping the map current without turning it into a news feed. Some disputed claims will become engineering ground; others will remain provisional longer than their slogans suggest. The rest of the book uses this orientation as a calibration tool, not as a closed theory. Part I turns that map into material choices: compute budgets, data mixtures, tokenizers, architectures, and training runs.

Comments

Log in to comment