Summary
This part stepped away from the left-to-right text spine and looked at generation when the object does not arrive as a natural token sequence. Diffusion, flow matching, non-autoregressive language modeling, speech, image generation, video, world models, and robotics all have to answer the same structural question: what order can the system invent so that learning, sampling, caching, and evaluation become possible?
That invented order is never free. A representation that helps training may make serving harder. Sharper sampling fidelity often arrives at the cost of multiplied latency, and a multimodal fusion scheme can obscure where the evidence came from. The central concern is therefore not media variety by itself, but the cost of translating a physical or perceptual object into something a model can handle.
The takeaway is that multimodal architecture is a discipline of constructed order. The open question is whether grounded experience and robotics data can scale with anything like the regularity of text. If they cannot, the bottleneck will move from model design to the collection, inheritance, and verification of physical-world data. Part III returns to behavior: once a model can generate through these orders, the next question is which behaviors should be encouraged, refused, ranked, or corrected.
Comments
Log in to comment