The AI Value Chain
A model call looks like a software service, but it is the visible end of a long industrial chain. The token that arrives in an editor or support desk has passed through silicon fabrication, accelerator allocation, power and cooling, training research, data rights, model serving, gateways, applications, and procurement. A useful economic map therefore cannot stop at "the model market." It has to ask where scarce inputs enter, where fixed costs are highest, where switching costs accumulate, and where value can still be competed away.
Capability is delivered through a chain
The chain is not linear in the product sense, but it is linear in its constraints. Scarce accelerators move training calendars. When data-center power is delayed, a cluster does not exist even after the purchase order has cleared, and a frontier model reachable only through an API hands its downstream applications that provider's price, latency, policy, and outage surface. Ambiguous data rights in a domain slow adaptation even when the model itself is good enough.
The concentration points follow the chain. Silicon and advanced packaging have large fixed costs, long lead times, and few suppliers. Cloud and power face regional bottlenecks and capacity commitments. Frontier training has a research fixed cost that rises with the capability frontier, as Chapter 76 showed through the Epoch AI cost curve (Cottier et al. 2024). APIs and gateways turn model capability into a priced digital service. Applications sit closer to the user and have lower capital cost, but they can gain switching cost through workflow, identity, data, permissions, and procurement.
The important correction is that open weights change only one link. They can reduce dependence on a model API, increase downstream experimentation, and make behind-frontier competition intense (Kapoor et al. 2024). They do not remove silicon concentration, cloud procurement, data rights, energy constraints, or the distribution advantage of the applications that already sit in the workflow.
Where concentration enters
Foundation models have two opposing economic forces. The frontier has a tendency toward concentration because the fixed costs of data, research, compute, talent, and evaluation are large, and because the best model can serve many markets at once. Vipra and Korinek describe this as a tendency toward natural monopoly for the most capable models, paired with a need for contestability and quality standards at the frontier (Vipra and Korinek 2023). Behind the frontier, competition can be much more intense: models are cheaper to train, open weights are available, distillation (training a small model to copy a larger one) works, and buyers can route between substitutes.
This split explains why the same market can look concentrated and competitive at the same time. A handful of organizations may train the most capable models, while hundreds of models compete for cheaper inference workloads, domain-specialized tasks, and local deployments. Du's token-pricing study reports a sharp decline in inference prices and falling concentration in inference services from 2020 to 2026, while also finding a reasoning premium for flagship models (Du 2026). The frontier rent remains where users pay for the last increments of quality, reasoning, reliability, tool use, or brand assurance.
The formal condition is simple enough to be useful:
Here is the fixed cost of entering the layer, is marginal cost, and is the customer's cost of leaving. Silicon carries a high . Cloud adds a regional on top of its own high , while a frontier API pairs a high upstream with sometimes high downstream, the latter built from prompts, evals, safety policy, fine-tunes, logging, and procurement. Applications enter with a lower , yet can accumulate high through the workflow and the data they gather.
Where the profit pools sit
The concentration condition predicts where the rent lands, and the measured profit pools confirm it. In 2026 the gross profit of the AI stack sits overwhelmingly at the bottom: roughly four fifths at the semiconductor layer, the rest split between infrastructure and applications. Capital flows the same direction. The five largest hyperscalers are on course to spend around $600 billion in 2026, some $450 billion of it on AI data centers and chips, money poured into the two layers that already hold more than nine tenths of the profit (Agrawal 2026).
| Layer | 2024 profit share | 2026 profit share | 2026 gross profit |
|---|---|---|---|
| Semiconductors | ~87% | ~79% | ~$225B |
| Infrastructure | ~9% | ~13% | ~$40B |
| Applications | ~4% | ~8% | ~$20B |
These are an investor's estimate rather than audited accounts, but the shape is the point (Agrawal 2026). It is the same , , condition reading out as dollars. Silicon carries the highest fixed cost and the lowest marginal cost at scale, so it holds the rent; an application, with low fixed cost and a substitute one routing call away, captures the least. The chip the model runs on earns more of a token's profit than the product the user actually pays for.
The trend is the part worth watching. The silicon share has given up about four points in two years, with infrastructure and applications splitting the gain. Agrawal reads this as the opening of an inversion: in earlier compute supercycles value migrated up from the foundational layer to the software built on it, and the cloud stack took on the order of fifteen years to move from hardware-dominated to software-dominated economics. At four points every two years, the application layer reaching the profit share that apps enjoy in cloud is a decade-out proposition, not a next-year one. The direction is up the stack; the clock is slow.
The open-model debate often treats openness as a binary answer to concentration. It is not. Open weights can reduce API lock-in and raise behind-frontier competition, but open weights do not automatically make a model "Open Source AI" under the OSI definition, and they do not replicate the power, capital, data, and distribution positions upstream and downstream (Open Source Initiative 2024; White et al. 2024). The open release changes the competitive pressure at the model artifact layer. It does not dissolve the value chain around it.
Vertical integration and the return loop
The strategic move in this chain is vertical integration. A chip vendor sells systems, invests in clouds, and exposes software libraries; a cloud signs capacity deals with model developers and bundles hosted APIs. A model developer builds consumer products and enterprise gateways, and at the far end an application vendor embeds models into the workflows where data is created. Each integration closes a loop:
- the upstream layer gets demand certainty and usage data;
- the downstream layer gets privileged access, better latency, or lower price;
- the customer receives a simpler bundle but loses some bargaining surface.
This is why transparency matters as market infrastructure, not just as academic virtue. If a buyer cannot see training data, training compute, incident history, usage monitoring, or downstream impact, it cannot price the risk of dependence. The 2025 Foundation Model Transparency Index reports that average developer transparency fell from 58 to 40 out of 100, with training data, training compute, and post-deployment usage among the most opaque categories (Wan et al. 2025). In a concentrated layer, opacity makes switching and regulatory oversight harder.
The value chain sends constraints upward and downward. Energy and capacity constraints from Chapter 68 determine what can be trained and served. Transparency and data-rights constraints from Chapter 79 determine which models can be bought by regulated customers. Adoption constraints from Chapter 78 determine which application layer captures value. A model market is therefore not decided by benchmark score alone; it is decided by which layer owns the binding constraint.
What to watch
For a practitioner, the useful question is not "who wins AI." It is which layer in your workload holds the scarce input:
- If the scarce input is compute, negotiate capacity and design for routing.
- If it is data rights, invest in provenance and contracts before adaptation.
- If it is model quality, keep a frontier option and measure switching cost.
- If it is workflow adoption, spend less time on model comparison and more on review loops, change management, and product fit.
Market structure is the economic shadow of the architecture. The stack in the earlier parts explained how capability is made. The value chain explains who can make it available, who can meter it, and where a team can still choose.
Further reading
- Vipra & Korinek, “Market Concentration Implications of Foundation Models” (natural-monopoly tendency at the frontier and intense behind-frontier competition), 2023. arXiv:2311.01550Vipra and Korinek argue that the most capable foundation models tend toward natural monopoly, while behind-frontier models can face intense competition, shaping antitrust and regulatory priorities.
- Kapoor et al., “On the Societal Impact of Open Foundation Models” (open foundation model benefits, risks, and marginal-risk framing), 2024. arXiv:2403.07918This position paper analyzes open-weight foundation models through their benefits, risks, and marginal risk relative to existing technologies, clarifying where evidence is still missing.
- Wan et al., “The 2025 Foundation Model Transparency Index” (annual transparency index for foundation model developers), 2025. arXiv:2512.10169The 2025 FMTI reports that average developer transparency fell from 58 to 40 out of 100, with training data, training compute, and post-deployment usage among the most opaque areas.
- Du, “Tiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in Large Language Model Inference Services” (token-pricing competition and concentration data), 2026. arXiv:2603.28576Du documents rapid token-price decline across model tiers and a fall in inference-market concentration, framing token prices as a distinct digital-goods market.
- Agrawal, “The Economics of Generative AI: Two Years Later” (gross-profit share by stack layer and the slow inversion thesis), 2026. apoorv03.comAgrawal estimates that in 2026 semiconductors still capture about 79 percent of AI gross profit against 13 percent for infrastructure and 8 percent for applications, down from roughly 87/9/4 two years earlier, and argues value will migrate up the stack as in past compute supercycles, but slowly.
Comments
Log in to comment