Summary
The infrastructure part went below the model. It began with accelerator bandwidth, then the software layer in between: autodiff and the frameworks that industrialized it, and the compiler and kernel layer where bytes moved, not FLOPs performed, set the price and one vendor's software sediment forms the industry's deepest moat. From there it covered cluster orchestration, data infrastructure, silicon, power, and failure at scale.
The real limit is rarely the one on the model card. It may be HBM, a network tier, a power interconnect, an export rule, a checkpoint system, or a failure distribution. The limit is a moving set of physical and operational constraints, and a design that ignores them is a design that will not run.
The open question this part leaves is one of supply. Whether advanced packaging, HBM, and grid interconnection can grow on the schedule the training roadmaps assume is not settled by chip design. It is settled by export rules, fab capacity, and a utility queue measured in years. What the hardware cannot supply at any schedule is a limit of another kind: more HBM does not manufacture training data, a faster interconnect does not report how capable the resulting model is, and neither one makes a produced claim cheaper to check. Part X takes up those limits, which sit above the machine rather than below it.
Comments
Log in to comment