AI Infra
0%
Part IX · Chapter 71

The Capability Horizon and Its Measurement

AuthorChangkun Ou
Reading time~12 min

Every layer of the book has asked whether a system can do the task, at what cost, and whether it can be trusted. The last chapter of this infrastructure arc asks the question the whole stack eventually has to answer: how capable is the frontier, and how would we know? In 2026 the frontier is no longer a leaderboard number but a moving horizon. The most useful measure of that horizon also masks its sharpest limit, and the instruments saturate faster than they can be built. The field's deepest disagreement, whether the recent acceleration is durable progress or a compute-bought artifact, turns less on another model result than on better measurement. The headline number in this chapter will be stale by the time it is read, and that staleness is the point.

2026-06-21T21:25:29.927050 image/svg+xml Matplotlib v3.11.0, https://matplotlib.org/
Figure 71.1. Schematic of the capability horizon. Short tasks saturate first, while longer tasks move later and make progress visible only when measurement follows the horizon outward. Idealized curves, not measured data.

Frontier as a horizon, not a score

For years progress was a column of benchmark percentages. The trouble is that the percentages saturate: once the best models pass the high eighties and nineties on a test, as they did on MMLU (Massive Multitask Language Understanding, a broad knowledge exam) and GPQA (graduate-level, Google-proof science questions), the test stops distinguishing them, and a saturated test measures nothing at the frontier. The reframing that organizes this chapter replaces the score with a horizon: the length of human task a model can complete autonomously. Measured this way, the frontier is a duration, and the duration has been climbing fast. The reference effort put a leading model near a five-hour task-completion horizon at the start of 2026 and, weeks later, near fifteen hours, with the doubling time of the trend dropping from roughly seven months to a few (METR 2026). By the time this sentence is read the number will be larger and a different model will hold it. The metric is valuable precisely because it is continuous and deployment-flavored rather than saturating, but it is also where the chapter's central tension lives.

score saturating scores (MMLU, GPQA near ceiling) horizon task-completion horizon ~5h → ~15h at 50%, early 2026 doubling ~7mo → a few score->horizon replaced by gap but 80% reliability horizon ≈ 5× shorter (the useful threshold) horizon->gap
Figure 71.2. Frontier reframed as a horizon: the length of task a model completes autonomously, climbing fast while the doubling time shortens. The headline is the 50%-reliability horizon; the economically useful 80% horizon is roughly five times shorter, and the gap is the binding constraint.

The number that masks the limit

The horizon is reported at a success rate, and the rate chosen flatters the result. The headline figure is the length of task a model finishes half the time. Real work needs more than a fifty-percent success rate, and the horizon at a high-reliability threshold, finishing the task eight times in ten, is roughly five times shorter than the headline (Kwa et al. 2025). A model with a fifteen-hour fifty-percent horizon has perhaps a three-hour eighty-percent horizon, and it is the second number that bounds what the model can be trusted to own. This is the binding constraint of the chapter: the gap between the task a model can sometimes finish and the task it can reliably finish, and the second lags the first by a wide and persistent margin.

The same compounding logic that broke long agent runs in Chapter 69 is at work here: per-step reliability compounds badly over a long task's many steps, so even high single-step competence yields a low chance of finishing a long task cleanly. The horizon metric measures the symptom; the compounding is the cause.

Constraint arrow

The fifty-versus-eighty-percent reliability gap, a property of per-step success rates at the model layer, caps the economically useful horizon at roughly a fifth of the headline. Reliability at the bottom dictates which tasks the autonomy layer above can own (cross-ref Chapter 52). And the slope of the whole trend is set by a lower layer too: the recent speedup came largely from reinforcement learning and test-time compute, so the frontier metric's pace is dictated by a post-training compute choice rather than by pretraining scale. Pretraining is the one large pass that builds a base model from raw data; post-training is the later, targeted work, including reinforcement learning and extra inference-time compute, that sharpens it (cross-ref Chapter 30, Chapter 28).

Tying capability to value

If the horizon is one attempt to make capability mean something, economic evaluation is the other. Rather than quiz questions, an economic benchmark grades real work products, deliverables drawn from many occupations and judged blind by experts. On the most direct such effort, the opening result in September 2025 had the best model, Claude Opus 4.1, producing deliverables rated as good as or better than the human expert's on just under half of the tasks, 47.6 percent, at roughly a hundredth of the time and cost (OpenAI 2025). "Just under half" read at the time as approaching parity, not at it. That reading did not last a year: by 2026 the best model won or tied the blind expert comparisons on roughly seven tasks in ten (Epoch AI 2026), so on the benchmark's own terms the crossing has happened, and the open question has moved from whether models reach graded parity to what a won comparison is worth outside the grading room. A parallel line asks whether models can do the research that builds the next model, and finds a revealing shape: agents beat human experts at a short time budget and lose to them at a long one (Wijk and others 2025), which is the horizon limit again, recovery and long-range planning failing where short competence succeeds.

Whether any of this translates into measured productivity is its own contested question, and the cleanest study is a caution. A controlled trial in early 2025 found experienced developers were about 19 percent slower with AI assistance while believing they were faster (METR 2025). That result is a snapshot of early-2025 tools, not a law: the same group later reported that uplift appears positive with 2026 tools, while warning that selection effects make the new signal weak. The careful statement is that capability and measured productivity are not the same quantity, that self-reports overstate, and that the transfer from benchmark to workplace is real but smaller and slower than the headlines imply.

The instruments saturate

The reframing toward horizons and economic value was forced by saturation, and saturation is now chasing the new instruments too. Purpose-built hard benchmarks have short shelf lives: a deliberately difficult expert exam (Humanity's Last Exam, HLE) climbed from the mid-thirties past fifty percent within months of release (Center for AI Safety and Scale AI 2026), and an abstraction-and-reasoning challenge (ARC-AGI-2) designed to resist memorization went from essentially zero past the average human score to the mid-eighties in little over a year (Chollet et al. 2025; ARC Prize Foundation 2026). Even the horizon suite that anchors this chapter is now near saturation, which is why its latest headline carries a confidence interval spanning most of an order of magnitude. The instrument that was built to outlast the others is hitting the same wall, and its maintainers are building a successor. This is the chapter's central irony in one fact: the measurement of the frontier ages as fast as the frontier.

Figure 71.3. A benchmark's top score climbing from near zero to saturation over the months after it is released. Drag the steepness to see how fast a hard benchmark gets solved: the faster it saturates, the sooner the next instrument has to be built.
saturation mmlu MMLU / GPQA saturated (>88%, ~94%) hard HLE / ARC-AGI-2 built hard, climbing fast mmlu->hard forced horizon task-horizon suite now near saturation hard->horizon forced next successor instruments (under construction) horizon->next forces
Figure 71.4. The instruments saturate faster than they can be built. Each generation of benchmark is created to resist saturation, climbs steeply as models exhaust it, and forces the next instrument. The horizon suite, built to outlast scores, is itself now near saturation.

The response is not only harder tests but a rebuilding of measurement itself. A construct-validity movement audits whether a benchmark measures any well-defined ability at all, finding that many do not and so overclaim generality. A general-scales program rebuilds evaluation around ability profiles and instance-level prediction, aiming to make a score explanatory rather than a bare number (Hernández-Orallo and others 2026; Construct Validity Review 2025). Others turn the lens on the horizon metric directly, asking whether its human-time baselines can be inferred more cheaply and whether the trend reflects a real capability or an artifact of how the tasks were chosen. Alongside this is the harder evidence that benchmarks can be gamed: several agentic suites have been driven to near-perfect scores without actually solving the tasks, which turns "benchmark theater" from a complaint into a demonstrated failure mode.

Is the acceleration real?

This brings the infrastructure arc to its hardest open question, which it should hold open rather than pretend to resolve. The horizon is climbing and the doubling time has shortened. One camp reads this as durable acceleration, a frontier compounding toward general capability and, eventually, toward systems that improve their own successors; the most sober version of that projection puts substantial automation of AI research in the early 2030s, under aggressive assumptions about compute and software efficiency, and labels itself a projection rather than a result, since no published system yet materially improves its own successor in a closed loop (METR 2026). The other camp reads the speedup as compute-bought: the recent gains came largely from reinforcement learning and test-time compute, the confidence intervals are wide, and the doubling may revert toward its older pace once that compute becomes constrained. The strongest statements of the plateau case argue that the era of progress-by-scale is ending and that the frontier will be defined by the cost of adaptability rather than by model size (Hooker 2026); the strongest continuation case holds that scaling is shifting axes, not stopping.

What's contested
  • Durable acceleration or compute-bought artifact? One camp sees the shortened doubling time as a real, continuing trend. The other, including the measurers' own caveats, sees a one-time gain from reinforcement learning and test-time compute that reverts once compute binds. The confidence intervals are wide enough to fit both.
  • Has scaling hit a wall that bounds the frontier? A peak-data camp argues pretraining-by-scale is ending and adaptability becomes the binding cost. A continuation camp argues the growth simply moved to post-training and inference-time compute. The mechanics belong to Chapter 5; what is contested here is whether the measured capability trend continues.
  • Do benchmarks measure general capability at all? A construct-invalidity camp says most lack a defined construct and overclaim. A reformist camp wants to rebuild evaluation around ability profiles. A pragmatist camp treats leaderboards as screening only and trusts layered human review on one's own tasks.
  • Real evidence or benchmark theater? Saturation, contamination, and demonstrated exploitability argue the headline scores are marketing. Purpose-built, blind-graded, contamination-controlled evals argue they still differentiate and still predict deployment better than old quizzes.
  • Does capability become productivity? Perceived uplift is large; one controlled trial found a slowdown; the later signal is positive but weak. The transfer is real, smaller than claimed, and hard to measure.

The horizon as a measuring instrument

At the horizon layer, the closing object is the instrument rather than the model. Capability, rigorously measured, is a duration that is genuinely rising and genuinely shorter than the headline once reliability is demanded. Efficiency now shapes the trend itself: the acceleration was bought with inference-time and post-training compute, so the slope of progress is a spending decision as much as a scientific one, which loops back to the compute and power ceilings that opened this part. Trust is the unresolved requirement. The field cannot yet certify that its numbers mean what they claim, because the instruments saturate, can be gamed, and do not cleanly transfer to work; an enterprise that needs to trust a model still falls back on layered human review of its own tasks.

The frontier in 2026 is therefore not a number. It is a set of moving limits, many below the chip and many above the model. Measurement is one of those upper limits, but not the last one. Chapter 72 takes the next step: if a model can produce a result that no ordinary reviewer can quickly check, the frontier has moved from capability to acceptance.

Further reading

  • METR, “Time Horizon 1.1” (the horizon metric and the 50), 2026. metr.org
    METR releases Time Horizon 1.1, an updated benchmark measuring AI agent task-completion time horizons using more tasks and a new eval infrastructure.
  • Kwa et al., “Measuring AI Ability to Complete Long Tasks” (the horizon metric and the 50), 2025. metr.org
    METR proposes measuring AI progress by the length of tasks agents can autonomously complete, finding this horizon has doubled roughly every 7 months over the past 6 years.
  • OpenAI, “GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks” (economic task completion: deliverables judged against human experts), 2025. arXiv:2510.04374
    GDPval is a benchmark of 1,320 economically grounded tasks across 44 occupations and 9 U.S. GDP sectors, showing frontier models approaching expert parity via pairwise human comparison.
  • Wijk & others, “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts” (AI research-and-development automation: agents win at short budgets, lose at long ones), 2025. arXiv:2411.15114
    RE-Bench is a seven-task ML research engineering benchmark comparing frontier AI agents to 61 human experts, finding agents score 4x higher at 2-hour budgets but humans outperform them at 8-hour and longer budgets.
  • Center for AI Safety and Scale AI, “Humanity's Last Exam” (saturation-resistant frontier benchmarks that still climbed fast), 2026. agi.safe.ai
    Humanity's Last Exam (HLE) is a 2,500-question benchmark of expert-level problems across diverse fields, designed to be among the hardest tests for frontier AI models.
  • Chollet et al., “ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems” (saturation-resistant frontier benchmarks that still climbed fast), 2025. arXiv:2505.11831
    ARC-AGI-2 is an upgraded abstract reasoning benchmark with harder, less brute-forcible tasks and large-scale first-party human baselines, designed to better measure progress toward general AI beyond ARC-AGI-1.
  • Hernández-Orallo & others, “General Scales Unlock AI Evaluation with Explanatory and Predictive Power” (the measurement-science overhaul), 2026. arXiv:2503.06378
    This paper proposes 18 general demand-level rubrics (ADeLe) to annotate LLM benchmark instances, enabling both explanatory ability profiles and instance-level performance prediction in- and out-of-distribution.
  • Construct Validity Review, “Measuring What Matters: Construct Validity in Large Language Model Benchmarks” (the measurement-science overhaul), 2025. openreview.net
  • METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (the 19), 2025. arXiv:2507.09089
    A randomized controlled trial with 16 experienced open-source developers found that early-2025 AI tools (Cursor Pro with Claude 3.5/3.7 Sonnet) increased task completion time by 19
  • Hooker, “On the Slow Death of Scaling” (the plateau and self-improvement-projection debates), 2026. papers.ssrn.com
  • METR, “A Simpler AI Timelines Model Predicts 99% AI R&D Automation in  2032” (the plateau and self-improvement-projection debates), 2026. metr.org
    METR presents an 8-parameter model for forecasting AI timelines, predicting roughly 99
  • Epoch AI, “GDPval” (the 2026 win-or-tie rate against expert graders), 2026. epoch.ai
    Epoch AI's tracking of the GDPval benchmark, on which the best 2026 model wins or ties against human experts on roughly 71
  • ARC Prize Foundation, “ARC-AGI Leaderboard” (ARC-AGI-2 top scores as of mid-2026), 2026. arcprize.org
    The live ARC-AGI leaderboard, showing ARC-AGI-2 top scores in the mid-eighties by June 2026, well past the average human score.

Comments

Log in to comment