AI Infra
0%
Part XI · Chapter 78

Adoption and Productivity

AuthorChangkun Ou
Reading time~17 min

Capability is potential. Productivity is a measured change in accepted work. A model can draft faster while making a reviewer slower, raise one person's output while congesting a shared queue, or improve a benchmark without changing anything an organization ships. Adoption therefore needs a counterfactual: what would the same eligible work have cost and produced under the policy it replaces?

Chapter 77 asked whether a buyer has credible alternatives. This chapter asks whether using one of them changes work enough to matter. A defensible claim names the accepted unit of work, task population, worker population, workflow, tool version, review policy, and time window. Change any one of those fields and the old productivity estimate may no longer apply.

Freeze the unit before measuring the tool

Start with a completed workflow unit, not a job title or a model call. A support unit might be a ticket accepted as resolved after a fixed return window. A software unit might be a merged change that passes review, tests, and a delayed defect window. A policy unit might be an adjudicated memo whose claims satisfy a locked rubric. Attempts that are abandoned, escalated, or rejected stay in the denominator.

This task-level boundary matters. Occupational exposure studies estimate where language models could reduce task time at unchanged quality; they do not measure adoption, achieved productivity, job displacement, or return on investment (Eloundou et al. 2024). Jobs combine many tasks with coordination, accountability, and work that remains outside the tool. Aggregate only after the workflow units are measured.

Record the evaluation contract before the pilot begins:

Field What must be frozen Why it changes the estimate
Eligible unit trigger, required inputs, completion state, exclusions prevents easy work or abandoned work from silently leaving the denominator
Population task families, workers or teams, experience strata, locations bounds who and what the estimate describes
Treatment model, interface, retrieval, prompt library, tools, training makes “AI access” a reproducible intervention
Review policy reviewer role, rubric, escalation, defect window includes the labor and quality boundary around generation
Baseline current workflow, staffing, software, and service levels identifies the actual counterfactual
Horizon onboarding, stabilization, and steady-state periods separates learning from a short novelty effect

Productivity is not one proxy (Organisation for Economic Co-operation and Development 2001). Keep a decision vector containing accepted throughput, active labor time, end-to-end cycle time, first-pass quality, review time, rework, escalation, severity-weighted defects, incidents, customer outcomes, and worker experience. For assignment arm zz, three useful summaries are

Pz=AzHz,Fz=Az(1)Nz,Dz=s=1SwsnzsNz.P_z = \frac{A_z}{H_z}, \qquad F_z = \frac{A_z^{(1)}}{N_z}, \qquad D_z = \frac{\sum_{s=1}^{S} w_s n_{zs}}{N_z}.

Here, z=0z=0 denotes the baseline arm and z=1z=1 the assisted arm. PzP_z is accepted throughput; AzA_z is the number of accepted units; and HzH_z is total paid labor hours, including generation, checking, correction, escalation, and coordination. FzF_z is first-pass yield, where Az(1)A_z^{(1)} counts units accepted without correction and NzN_z counts every assigned eligible unit. DzD_z is the severity-weighted defect rate; ss indexes one of SS predeclared severity classes, wsw_s is its fixed weight, and nzsn_{zs} is the number of class-ss defects in arm zz. Report raw defect counts beside DzD_z so a chosen weight cannot hide harm.

Prompt count, tokens generated, logins, and satisfaction can explain a mechanism. They are not productivity. Neither is faster completion if acceptance quality, rework, or delayed defects deteriorate.

Estimate the effect of access, not enthusiasm

When feasible, randomly assign access within prespecified strata such as task family, baseline output, and worker experience. Here, ii indexes the unit of randomization, Zi=1Z_i=1 indicate assignment to access, Zi=0Z_i=0 indicate assignment to the baseline, and YiY_i be the locked primary outcome for that unit. With n1n_1 assisted assignments and n0n_0 baseline assignments, the sample intent-to-treat effect is

τ^ITT=1n1i:Zi=1Yi1n0i:Zi=0Yi.\widehat{\tau}_{\mathrm{ITT}} = \frac{1}{n_1}\sum_{i:Z_i=1}Y_i - \frac{1}{n_0}\sum_{i:Z_i=0}Y_i.

The intent-to-treat estimate measures the policy effect of offering this tool under these conditions. It retains people who were assigned access but did not use it. Actual use is a post-assignment choice. Comparing voluntary users with nonusers confounds the tool with motivation, task selection, and expected benefit; calling that difference causal creates selection bias. Noncompliance can also be reported through take-up and, with the required instrumental-variable assumptions, a complier effect, but that is not the average effect for every user (Angrist et al. 1996).

The unit of randomization must match the interference boundary. If colleagues share a queue, templates, reviewers, or meeting load, individual treatment can spill over into the control arm. Cluster-randomize the team or workflow when those interactions are material, and analyze at that level. A network or saturation design is needed when direct and spillover effects are separate decisions of interest (Hudgens and Halloran 2008).

Predeclare a small number of useful strata, such as task family, experience, error severity, and workflow maturity. Report effect sizes, uncertainty intervals, and subgroup counts. “Significant here but not there” is not proof that two subgroup effects differ. Keep the original random assignment through onboarding and learning periods, and version every change to the model, interface, prompt library, staffing, and task mix.

Put value and cost in the same ledger

Operational metrics should remain separate until the organization has a defensible conversion rule. When a financial decision is required, compare both arms over one accounting horizon HH:

NBH=(R1R0)+(C0C1)CfixedE[L1L0].NB_H = (R_1-R_0) + (C_0-C_1) - C_{\mathrm{fixed}} - \mathbb{E}[L_1-L_0].

NBHNB_H is incremental net benefit over horizon HH. RzR_z is realized revenue or contribution value in arm zz; CzC_z is observed variable operating cost, including labor, model, review, rework, escalation, and coordination; and CfixedC_{\mathrm{fixed}} is incremental integration, security, legal, procurement, and training cost. LzL_z is loss from errors and incidents, so E[L1L0]\mathbb{E}[L_1-L_0] is the change in expected loss. Every term uses the same currency and accounting horizon. Arm 0 is the baseline arm and arm 1 the assisted arm.

Avoid two common accounting errors. First, an hour saved is capacity, not cash, unless the organization can redeploy it, avoid hiring, increase accepted output, or reduce paid hours. State that conversion rule. Second, do not count a quality improvement once as revenue and again as avoided rework. If quality or a rare catastrophic loss cannot be credibly priced, keep it as a separate guardrail and stress scenario.

For a steady-state unit-cost decision, Cfixed/NHC_{\mathrm{fixed}}/N_H is the amortized fixed cost per accepted unit, where NHN_H is planned accepted volume over HH. Keep observed pilot cost separate from forecast volume. Report an uncertainty interval for NBHNB_H and sensitivity to labor value, volume, incident probability, and the defect window rather than presenting one precise return.

The explorer below uses illustrative normalized value units, not percentages, prices, evidence, or a forecast. Its purpose is to expose which assumptions can reverse the decision. A real pilot must populate the ledger from the same accounting period in both arms.

Figure 78.1. An illustrative sensitivity balance in normalized value units. Move each benefit and cost independently; the values are assumptions, not measured productivity.

Why a useful tool can look unproductive at first

The gap between capability and measured output has a history. The productivity J-curve describes how a general-purpose technology can require complementary investment before its benefits appear in measured output. Organizations first build intangible capital: new processes, data, training, management routines, and software. During that period, the investment is costly while its future services are poorly measured (Brynjolfsson et al. 2021).

Generative AI does not remove that mechanism. Business-process redesign may change who drafts, who reviews, which queue owns an exception, and how accepted work reaches a customer. A separate chat window can improve individual drafting without changing a shared approval path. The resulting measurement lag is not a reason to assume that benefits will arrive later; it is a reason to measure complementary investment, adoption time, and the steady-state workflow instead of extrapolating from a demonstration.

Read every empirical result with its boundary

Early studies answer different questions and do not combine into one universal effect size.

Noy and Zhang randomly assigned 453 college-educated professionals to complete incentivized, midlevel writing tasks with or without ChatGPT. Access reduced completion time by 40 percent and raised output quality by 18 percent on those short tasks (Noy and Zhang 2023). The experiment establishes a bounded task effect, not a change in team output, employment, or firm profit.

Dell'Acqua and colleagues ran a preregistered experiment with 758 consultants at one firm. The study used two task bundles: 18 product, analysis, and writing tasks selected to be inside GPT-4's June 2023 capability boundary, and a business case containing one deliberately selected task outside it. In the inside bundle, participants with AI completed 12.2 percent more tasks and worked 25.1 percent faster, with higher rated quality. On the selected outside task, the combined AI groups were 19 percentage points less likely to be correct (Dell'Acqua et al. 2026). This contrast motivates a jagged frontier, but it does not define a permanent frontier for other models, workers, or workflows. The legacy synthetic curve has therefore been removed.

Brynjolfsson, Li, and Raymond observed the staggered managerial rollout of a real-time assistant across 5,172 customer-support agents, three million chats, and a single firm with contractors. Their main difference-in-differences estimate was a 15 percent increase in resolved issues per hour, with larger gains for less experienced and lower-skill agents. The highest-skill group saw little throughput gain and small quality declines on some measures (Brynjolfsson et al. 2025). This is strong field evidence for that assistant and setting, not a large randomized estimate across occupations.

Dillon and colleagues randomized integrated tool access among 7,137 workers at 66 opt-in firms. During months four through six, assignment reduced email-session time by 1.4 fewer hours per week; the instrumental-variable estimate for compliers was 2.0 fewer hours. The study detected no significant average change in meeting time, Word time, completed documents, or several coordination measures (Dillon et al. 2026). Telemetry showed a time-allocation effect, not output quality or organizational return. It also illustrates why individual drafting speed need not move a coordinated routine.

Becker and colleagues randomized whether 16 developers could use early-2025 tools on 246 tasks drawn from issues in mature open-source projects they knew well. AI allowance increased completion time by 19 percent, with a 95 percent interval from 2 to 39 percent slower, even though participants believed the tools had helped (Becker et al. 2025). The result concerns short brownfield issues, codebase experts, and February-to-June 2025 tools. It does not estimate current tools, greenfield development, long-run code quality, or team throughput. METR later changed its follow-up design because growing AI dependence made a representative no-AI control group difficult to recruit, a selection problem rather than a clean reversal of the first estimate (METR 2026).

Broader outcomes can move less than task behavior. A 2026 revision linking surveys in Denmark with administrative records found rapid chatbot use and task reorganization, including new oversight and integration work, but no detectable average effect on earnings or recorded hours. Its estimates ruled out effects larger than 2 percent at worker and workplace levels two years after ChatGPT's launch (Humlum and Vestergaard 2025). That result does not deny local time savings. It shows that saved task time, changed work, hours, earnings, and aggregate productivity are different outcomes.

Separate adoption, usage, capability, and realized value

Usage logs locate requests. They do not show whether an output was correct, accepted, or useful. Anthropic's initial Economic Index mapped more than four million Claude conversations in total; its central occupational analysis used self-selected Free and Pro conversations and excluded enterprise and API traffic. Software-development and writing requests dominated that product sample, but the study did not observe the user's job, final output, or productivity (Handa et al. 2025).

GDPval asks a different question. It gives a model an isolated deliverable, supplied context, and a one-shot task from one of 44 knowledge-work occupations. Expert graders found the strongest 2025 models approaching expert deliverable quality on the public 220-task gold set (Patwardhan et al. 2025). The benchmark excludes iterative client feedback, organizational context, integration, and most human oversight. Neither establishes causal productivity or an organizational return.

Adoption surveys supply another layer. A nationally representative U.S. Census Bureau supplement covering November 2025 through January 2026 found that 18 percent of U.S. firms reported AI use in a business function. Among adopting firms, 57 percent used it in no more than three functions (Bonney et al. 2026). Those self-reports measure adoption and breadth, not causal output gains; the paper's performance relationships are correlations.

Evidence What it can establish What remains unmeasured
Capability benchmark performance under a fixed task and grader adoption, workflow cost, spillovers, profit
Product usage log requests observed on that product and sample accepted output, counterfactual, causal effect
Adoption survey reported use and organizational breadth actual use quality, net benefit, mechanism
Controlled task experiment causal effect for its task, people, and treatment untested workflows and equilibrium effects
Field experiment or rollout operating effect under a real deployment generalization beyond its firm, period, and policy
Administrative outcome study downstream hours, earnings, or employment local quality and process mechanisms without linked logs
What's contested

There is no stable category called “AI-assisted work.” Tool capability changes, workers learn, task mix shifts, and organizations redesign around successful uses. A positive average can hide harm in a high-severity task; a null firm-level effect can coexist with valuable local capacity; and an early slowdown can be an investment or a genuine failure. The claim must stay attached to its population, treatment version, outcome, review boundary, and date. Heterogeneity is a result to estimate, not a license to explain any outcome after seeing it.

Run an adoption pilot as a decision process

A pilot should finish with evidence for a bounded policy, not a sentiment score:

  1. Freeze the unit. Declare eligible work, the acceptance rule, defect window, task strata, worker population, and current baseline.
  2. Register the design. Lock the unit of randomization, primary outcome, minimum worthwhile effect, sample and duration, subgroup plan, missing-data rule, interim looks, harm boundaries, and rollback trigger.
  3. Instrument both arms. Capture assignment, actual use, active labor, review, rework, escalation, coordination, model cost, failures, and versioned workflow events for every eligible unit.
  4. Apply the acceptance rule. Use a condition-blind review where feasible, preserve failures in the denominator, and wait through the delayed-defect window.
  5. Estimate effects. Report arm totals, take-up, intent-to-treat effects, uncertainty intervals, predeclared heterogeneity, spillovers, and learning by event time. Chapter 48 supplies the interval and stopping discipline.
  6. Price the ledger. Convert only defensible benefits and costs, state the capacity-to-cash assumption, retain unpriced quality and safety guardrails, and stress rare losses.
  7. Make a bounded decision. Go only when the lower uncertainty bound clears the hurdle and every guardrail passes; iterate when uptake or workflow friction prevents a conclusion; stop and rollback at a predeclared harm boundary.
adoption_pilot U Freeze work unit D Register design U->D I Instrument both arms D->I A Apply acceptance rule I->A E Estimate effects A->E P Price common-unit ledger E->P B Go | iterate | rollback P->B
Figure 78.2. A productivity claim moves from a frozen work unit to a bounded deployment decision.

The evidence record should retain the protocol and its amendments, assignment table, task and worker strata, treatment versions, raw arm totals, missingness, acceptance decisions, defect observations, cost rules, estimates and intervals, guardrail outcomes, and final decision. NIST's generative-AI risk profile likewise connects predeployment testing to documented go/no-go decisions, incident response, and postdeployment monitoring (National Institute of Standards and Technology 2024).

Constraint arrow

Chapter 77 determines whether the pilot has a credible alternative; Chapter 76 supplies model and infrastructure costs; and Chapter 48 supplies the inference contract. The next chapter, Chapter 79, adds the rights and compliance conditions that can make an otherwise productive workflow unusable. The economic unit remains one accepted and accountable work result, not one token or one model call.

Capability is potential. Productivity is a measured change in accepted work.

Further reading

  • Organisation for Economic Co-operation and Development, “Measuring Productivity: OECD Manual” (definitions and measurement guidance for productivity levels and growth), 2001. doi.org
    The manual explains why productivity must relate a specified output measure to specified inputs and documents the choices needed for valid comparisons.
  • Eloundou et al., “GPTs are GPTs: Labor Market Impact Potential of LLMs” (task-exposure estimates rather than observed adoption or productivity), 2024. doi.org
    The paper estimates how much occupational task time could be affected by language models while holding task quality constant; it does not measure realized workplace effects.
  • Angrist et al., “Identification of Causal Effects Using Instrumental Variables” (the potential-outcomes interpretation of instrumental-variable estimates), 1996. nber.org
    The paper states the assumptions under which randomized assignment can identify an average causal effect for compliers when treatment take-up is incomplete.
  • Hudgens & Halloran, “Toward Causal Inference with Interference” (causal designs for settings where one unit's treatment can affect another), 2008. pmc.ncbi.nlm.nih.gov
    The paper defines direct, indirect, total, and overall causal effects when treatments spill over within groups.
  • Brynjolfsson et al., “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies” (why complementary intangible investment can precede measured productivity gains), 2021. doi.org
    The authors model and document how costly investment in processes and other intangible capital can initially conceal the output gains from a general-purpose technology.
  • Noy & Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence” (randomized evidence from bounded professional writing tasks), 2023. doi.org
    In an experiment with 453 professionals, ChatGPT access reduced completion time and improved rated quality on short, incentivized writing tasks.
  • Dell'Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality” (a 758-consultant experiment on tasks inside and outside a selected capability boundary), 2026. doi.org
    Consultants gained speed and quality on selected tasks inside GPT-4's capability boundary but became less accurate on a deliberately selected task outside it.
  • Brynjolfsson et al., “Generative AI at Work” (a staggered real-world rollout to 5,172 customer-support agents), 2025. doi.org
    A conversational assistant increased resolved issues per hour at one customer-support operation, with larger gains among less experienced and lower-skill agents.
  • Dillon et al., “Shifting Work Patterns with Generative AI” (forthcoming; a six-month randomized field experiment with 7,137 eligible workers), 2026. aeaweb.org
    Integrated AI access reduced email-session time but produced no detectable average change in meetings, Word time, completed documents, or several coordination measures.
  • Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (a randomized trial on familiar issues in mature open-source repositories), 2025. arXiv:2507.09089
    Sixteen experienced developers took 19 percent longer with early-2025 AI tools on 246 issues, despite expecting and perceiving speed gains.
  • METR, “We Are Changing Our Developer Productivity Experiment Design” (an update on selection problems in a follow-up developer study), 2026. metr.org
    METR explains why increasing AI use among developers made its planned no-AI comparison group unrepresentative and required a different experimental design.
  • Humlum & Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI” (revised 2026; Danish surveys linked to administrative labor-market records), 2025. nber.org
    The linked Danish evidence finds rapid chatbot use and task reorganization but no detectable average effect on recorded hours or earnings two years after ChatGPT's launch.
  • Handa et al., “Which Economic Tasks Are Performed with AI? Evidence from Millions of Claude Conversations” (product-usage evidence mapped to occupational tasks), 2025. arXiv:2503.04761
    The study maps millions of sampled Claude conversations to occupational tasks, describing product usage without observing users' jobs, accepted outputs, or productivity.
  • Patwardhan et al., “GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks” (expert-graded isolated deliverables across 44 occupations), 2025. arXiv:2510.04374
    GDPval tests one-shot model deliverables on tasks from 44 occupations, measuring bounded capability rather than workplace adoption or net value.
  • Bonney et al., “The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks” (nationally representative evidence on reported firm adoption and breadth), 2026. census.gov
    The survey estimates firm AI adoption across business functions and shows that use remains narrow within most adopting firms; it does not identify causal productivity.
  • National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile” (risk-management actions for testing, deployment decisions, incidents, and monitoring), 2024. nvlpubs.nist.gov
    The NIST generative AI profile adapts the AI RMF to risks specific to generative systems, which this chapter translates into operating controls.

Comments

Log in to comment