Adoption and Productivity
Capability is potential. Productivity is a measured change in accepted work. A model can draft faster while making a reviewer slower, raise one person's output while congesting a shared queue, or improve a benchmark without changing anything an organization ships. Adoption therefore needs a counterfactual: what would the same eligible work have cost and produced under the policy it replaces?
Chapter 77 asked whether a buyer has credible alternatives. This chapter asks whether using one of them changes work enough to matter. A defensible claim names the accepted unit of work, task population, worker population, workflow, tool version, review policy, and time window. Change any one of those fields and the old productivity estimate may no longer apply.
Freeze the unit before measuring the tool
Start with a completed workflow unit, not a job title or a model call. A support unit might be a ticket accepted as resolved after a fixed return window. A software unit might be a merged change that passes review, tests, and a delayed defect window. A policy unit might be an adjudicated memo whose claims satisfy a locked rubric. Attempts that are abandoned, escalated, or rejected stay in the denominator.
This task-level boundary matters. Occupational exposure studies estimate where language models could reduce task time at unchanged quality; they do not measure adoption, achieved productivity, job displacement, or return on investment (Eloundou et al. 2024). Jobs combine many tasks with coordination, accountability, and work that remains outside the tool. Aggregate only after the workflow units are measured.
Record the evaluation contract before the pilot begins:
| Field | What must be frozen | Why it changes the estimate |
|---|---|---|
| Eligible unit | trigger, required inputs, completion state, exclusions | prevents easy work or abandoned work from silently leaving the denominator |
| Population | task families, workers or teams, experience strata, locations | bounds who and what the estimate describes |
| Treatment | model, interface, retrieval, prompt library, tools, training | makes “AI access” a reproducible intervention |
| Review policy | reviewer role, rubric, escalation, defect window | includes the labor and quality boundary around generation |
| Baseline | current workflow, staffing, software, and service levels | identifies the actual counterfactual |
| Horizon | onboarding, stabilization, and steady-state periods | separates learning from a short novelty effect |
Productivity is not one proxy (Organisation for Economic Co-operation and Development 2001). Keep a decision vector containing accepted throughput, active labor time, end-to-end cycle time, first-pass quality, review time, rework, escalation, severity-weighted defects, incidents, customer outcomes, and worker experience. For assignment arm , three useful summaries are
Here, denotes the baseline arm and the assisted arm. is accepted throughput; is the number of accepted units; and is total paid labor hours, including generation, checking, correction, escalation, and coordination. is first-pass yield, where counts units accepted without correction and counts every assigned eligible unit. is the severity-weighted defect rate; indexes one of predeclared severity classes, is its fixed weight, and is the number of class- defects in arm . Report raw defect counts beside so a chosen weight cannot hide harm.
Prompt count, tokens generated, logins, and satisfaction can explain a mechanism. They are not productivity. Neither is faster completion if acceptance quality, rework, or delayed defects deteriorate.
Estimate the effect of access, not enthusiasm
When feasible, randomly assign access within prespecified strata such as task family, baseline output, and worker experience. Here, indexes the unit of randomization, indicate assignment to access, indicate assignment to the baseline, and be the locked primary outcome for that unit. With assisted assignments and baseline assignments, the sample intent-to-treat effect is
The intent-to-treat estimate measures the policy effect of offering this tool under these conditions. It retains people who were assigned access but did not use it. Actual use is a post-assignment choice. Comparing voluntary users with nonusers confounds the tool with motivation, task selection, and expected benefit; calling that difference causal creates selection bias. Noncompliance can also be reported through take-up and, with the required instrumental-variable assumptions, a complier effect, but that is not the average effect for every user (Angrist et al. 1996).
The unit of randomization must match the interference boundary. If colleagues share a queue, templates, reviewers, or meeting load, individual treatment can spill over into the control arm. Cluster-randomize the team or workflow when those interactions are material, and analyze at that level. A network or saturation design is needed when direct and spillover effects are separate decisions of interest (Hudgens and Halloran 2008).
Predeclare a small number of useful strata, such as task family, experience, error severity, and workflow maturity. Report effect sizes, uncertainty intervals, and subgroup counts. “Significant here but not there” is not proof that two subgroup effects differ. Keep the original random assignment through onboarding and learning periods, and version every change to the model, interface, prompt library, staffing, and task mix.
Put value and cost in the same ledger
Operational metrics should remain separate until the organization has a defensible conversion rule. When a financial decision is required, compare both arms over one accounting horizon :
is incremental net benefit over horizon . is realized revenue or contribution value in arm ; is observed variable operating cost, including labor, model, review, rework, escalation, and coordination; and is incremental integration, security, legal, procurement, and training cost. is loss from errors and incidents, so is the change in expected loss. Every term uses the same currency and accounting horizon. Arm 0 is the baseline arm and arm 1 the assisted arm.
Avoid two common accounting errors. First, an hour saved is capacity, not cash, unless the organization can redeploy it, avoid hiring, increase accepted output, or reduce paid hours. State that conversion rule. Second, do not count a quality improvement once as revenue and again as avoided rework. If quality or a rare catastrophic loss cannot be credibly priced, keep it as a separate guardrail and stress scenario.
For a steady-state unit-cost decision, is the amortized fixed cost per accepted unit, where is planned accepted volume over . Keep observed pilot cost separate from forecast volume. Report an uncertainty interval for and sensitivity to labor value, volume, incident probability, and the defect window rather than presenting one precise return.
The explorer below uses illustrative normalized value units, not percentages, prices, evidence, or a forecast. Its purpose is to expose which assumptions can reverse the decision. A real pilot must populate the ledger from the same accounting period in both arms.
Why a useful tool can look unproductive at first
The gap between capability and measured output has a history. The productivity J-curve describes how a general-purpose technology can require complementary investment before its benefits appear in measured output. Organizations first build intangible capital: new processes, data, training, management routines, and software. During that period, the investment is costly while its future services are poorly measured (Brynjolfsson et al. 2021).
Generative AI does not remove that mechanism. Business-process redesign may change who drafts, who reviews, which queue owns an exception, and how accepted work reaches a customer. A separate chat window can improve individual drafting without changing a shared approval path. The resulting measurement lag is not a reason to assume that benefits will arrive later; it is a reason to measure complementary investment, adoption time, and the steady-state workflow instead of extrapolating from a demonstration.
Read every empirical result with its boundary
Early studies answer different questions and do not combine into one universal effect size.
Noy and Zhang randomly assigned 453 college-educated professionals to complete incentivized, midlevel writing tasks with or without ChatGPT. Access reduced completion time by 40 percent and raised output quality by 18 percent on those short tasks (Noy and Zhang 2023). The experiment establishes a bounded task effect, not a change in team output, employment, or firm profit.
Dell'Acqua and colleagues ran a preregistered experiment with 758 consultants at one firm. The study used two task bundles: 18 product, analysis, and writing tasks selected to be inside GPT-4's June 2023 capability boundary, and a business case containing one deliberately selected task outside it. In the inside bundle, participants with AI completed 12.2 percent more tasks and worked 25.1 percent faster, with higher rated quality. On the selected outside task, the combined AI groups were 19 percentage points less likely to be correct (Dell'Acqua et al. 2026). This contrast motivates a jagged frontier, but it does not define a permanent frontier for other models, workers, or workflows. The legacy synthetic curve has therefore been removed.
Brynjolfsson, Li, and Raymond observed the staggered managerial rollout of a real-time assistant across 5,172 customer-support agents, three million chats, and a single firm with contractors. Their main difference-in-differences estimate was a 15 percent increase in resolved issues per hour, with larger gains for less experienced and lower-skill agents. The highest-skill group saw little throughput gain and small quality declines on some measures (Brynjolfsson et al. 2025). This is strong field evidence for that assistant and setting, not a large randomized estimate across occupations.
Dillon and colleagues randomized integrated tool access among 7,137 workers at 66 opt-in firms. During months four through six, assignment reduced email-session time by 1.4 fewer hours per week; the instrumental-variable estimate for compliers was 2.0 fewer hours. The study detected no significant average change in meeting time, Word time, completed documents, or several coordination measures (Dillon et al. 2026). Telemetry showed a time-allocation effect, not output quality or organizational return. It also illustrates why individual drafting speed need not move a coordinated routine.
Becker and colleagues randomized whether 16 developers could use early-2025 tools on 246 tasks drawn from issues in mature open-source projects they knew well. AI allowance increased completion time by 19 percent, with a 95 percent interval from 2 to 39 percent slower, even though participants believed the tools had helped (Becker et al. 2025). The result concerns short brownfield issues, codebase experts, and February-to-June 2025 tools. It does not estimate current tools, greenfield development, long-run code quality, or team throughput. METR later changed its follow-up design because growing AI dependence made a representative no-AI control group difficult to recruit, a selection problem rather than a clean reversal of the first estimate (METR 2026).
Broader outcomes can move less than task behavior. A 2026 revision linking surveys in Denmark with administrative records found rapid chatbot use and task reorganization, including new oversight and integration work, but no detectable average effect on earnings or recorded hours. Its estimates ruled out effects larger than 2 percent at worker and workplace levels two years after ChatGPT's launch (Humlum and Vestergaard 2025). That result does not deny local time savings. It shows that saved task time, changed work, hours, earnings, and aggregate productivity are different outcomes.
Separate adoption, usage, capability, and realized value
Usage logs locate requests. They do not show whether an output was correct, accepted, or useful. Anthropic's initial Economic Index mapped more than four million Claude conversations in total; its central occupational analysis used self-selected Free and Pro conversations and excluded enterprise and API traffic. Software-development and writing requests dominated that product sample, but the study did not observe the user's job, final output, or productivity (Handa et al. 2025).
GDPval asks a different question. It gives a model an isolated deliverable, supplied context, and a one-shot task from one of 44 knowledge-work occupations. Expert graders found the strongest 2025 models approaching expert deliverable quality on the public 220-task gold set (Patwardhan et al. 2025). The benchmark excludes iterative client feedback, organizational context, integration, and most human oversight. Neither establishes causal productivity or an organizational return.
Adoption surveys supply another layer. A nationally representative U.S. Census Bureau supplement covering November 2025 through January 2026 found that 18 percent of U.S. firms reported AI use in a business function. Among adopting firms, 57 percent used it in no more than three functions (Bonney et al. 2026). Those self-reports measure adoption and breadth, not causal output gains; the paper's performance relationships are correlations.
| Evidence | What it can establish | What remains unmeasured |
|---|---|---|
| Capability benchmark | performance under a fixed task and grader | adoption, workflow cost, spillovers, profit |
| Product usage log | requests observed on that product and sample | accepted output, counterfactual, causal effect |
| Adoption survey | reported use and organizational breadth | actual use quality, net benefit, mechanism |
| Controlled task experiment | causal effect for its task, people, and treatment | untested workflows and equilibrium effects |
| Field experiment or rollout | operating effect under a real deployment | generalization beyond its firm, period, and policy |
| Administrative outcome study | downstream hours, earnings, or employment | local quality and process mechanisms without linked logs |
There is no stable category called “AI-assisted work.” Tool capability changes, workers learn, task mix shifts, and organizations redesign around successful uses. A positive average can hide harm in a high-severity task; a null firm-level effect can coexist with valuable local capacity; and an early slowdown can be an investment or a genuine failure. The claim must stay attached to its population, treatment version, outcome, review boundary, and date. Heterogeneity is a result to estimate, not a license to explain any outcome after seeing it.
Run an adoption pilot as a decision process
A pilot should finish with evidence for a bounded policy, not a sentiment score:
- Freeze the unit. Declare eligible work, the acceptance rule, defect window, task strata, worker population, and current baseline.
- Register the design. Lock the unit of randomization, primary outcome, minimum worthwhile effect, sample and duration, subgroup plan, missing-data rule, interim looks, harm boundaries, and rollback trigger.
- Instrument both arms. Capture assignment, actual use, active labor, review, rework, escalation, coordination, model cost, failures, and versioned workflow events for every eligible unit.
- Apply the acceptance rule. Use a condition-blind review where feasible, preserve failures in the denominator, and wait through the delayed-defect window.
- Estimate effects. Report arm totals, take-up, intent-to-treat effects, uncertainty intervals, predeclared heterogeneity, spillovers, and learning by event time. Chapter 48 supplies the interval and stopping discipline.
- Price the ledger. Convert only defensible benefits and costs, state the capacity-to-cash assumption, retain unpriced quality and safety guardrails, and stress rare losses.
- Make a bounded decision. Go only when the lower uncertainty bound clears the hurdle and every guardrail passes; iterate when uptake or workflow friction prevents a conclusion; stop and rollback at a predeclared harm boundary.
The evidence record should retain the protocol and its amendments, assignment table, task and worker strata, treatment versions, raw arm totals, missingness, acceptance decisions, defect observations, cost rules, estimates and intervals, guardrail outcomes, and final decision. NIST's generative-AI risk profile likewise connects predeployment testing to documented go/no-go decisions, incident response, and postdeployment monitoring (National Institute of Standards and Technology 2024).
Chapter 77 determines whether the pilot has a credible alternative; Chapter 76 supplies model and infrastructure costs; and Chapter 48 supplies the inference contract. The next chapter, Chapter 79, adds the rights and compliance conditions that can make an otherwise productive workflow unusable. The economic unit remains one accepted and accountable work result, not one token or one model call.
Capability is potential. Productivity is a measured change in accepted work.
Further reading
- Organisation for Economic Co-operation and Development, “Measuring Productivity: OECD Manual” (definitions and measurement guidance for productivity levels and growth), 2001. doi.orgThe manual explains why productivity must relate a specified output measure to specified inputs and documents the choices needed for valid comparisons.
- Eloundou et al., “GPTs are GPTs: Labor Market Impact Potential of LLMs” (task-exposure estimates rather than observed adoption or productivity), 2024. doi.orgThe paper estimates how much occupational task time could be affected by language models while holding task quality constant; it does not measure realized workplace effects.
- Angrist et al., “Identification of Causal Effects Using Instrumental Variables” (the potential-outcomes interpretation of instrumental-variable estimates), 1996. nber.orgThe paper states the assumptions under which randomized assignment can identify an average causal effect for compliers when treatment take-up is incomplete.
- Hudgens & Halloran, “Toward Causal Inference with Interference” (causal designs for settings where one unit's treatment can affect another), 2008. pmc.ncbi.nlm.nih.govThe paper defines direct, indirect, total, and overall causal effects when treatments spill over within groups.
- Brynjolfsson et al., “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies” (why complementary intangible investment can precede measured productivity gains), 2021. doi.orgThe authors model and document how costly investment in processes and other intangible capital can initially conceal the output gains from a general-purpose technology.
- Noy & Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence” (randomized evidence from bounded professional writing tasks), 2023. doi.orgIn an experiment with 453 professionals, ChatGPT access reduced completion time and improved rated quality on short, incentivized writing tasks.
- Dell'Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality” (a 758-consultant experiment on tasks inside and outside a selected capability boundary), 2026. doi.orgConsultants gained speed and quality on selected tasks inside GPT-4's capability boundary but became less accurate on a deliberately selected task outside it.
- Brynjolfsson et al., “Generative AI at Work” (a staggered real-world rollout to 5,172 customer-support agents), 2025. doi.orgA conversational assistant increased resolved issues per hour at one customer-support operation, with larger gains among less experienced and lower-skill agents.
- Dillon et al., “Shifting Work Patterns with Generative AI” (forthcoming; a six-month randomized field experiment with 7,137 eligible workers), 2026. aeaweb.orgIntegrated AI access reduced email-session time but produced no detectable average change in meetings, Word time, completed documents, or several coordination measures.
- Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (a randomized trial on familiar issues in mature open-source repositories), 2025. arXiv:2507.09089Sixteen experienced developers took 19 percent longer with early-2025 AI tools on 246 issues, despite expecting and perceiving speed gains.
- METR, “We Are Changing Our Developer Productivity Experiment Design” (an update on selection problems in a follow-up developer study), 2026. metr.orgMETR explains why increasing AI use among developers made its planned no-AI comparison group unrepresentative and required a different experimental design.
- Humlum & Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI” (revised 2026; Danish surveys linked to administrative labor-market records), 2025. nber.orgThe linked Danish evidence finds rapid chatbot use and task reorganization but no detectable average effect on recorded hours or earnings two years after ChatGPT's launch.
- Handa et al., “Which Economic Tasks Are Performed with AI? Evidence from Millions of Claude Conversations” (product-usage evidence mapped to occupational tasks), 2025. arXiv:2503.04761The study maps millions of sampled Claude conversations to occupational tasks, describing product usage without observing users' jobs, accepted outputs, or productivity.
- Patwardhan et al., “GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks” (expert-graded isolated deliverables across 44 occupations), 2025. arXiv:2510.04374GDPval tests one-shot model deliverables on tasks from 44 occupations, measuring bounded capability rather than workplace adoption or net value.
- Bonney et al., “The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks” (nationally representative evidence on reported firm adoption and breadth), 2026. census.govThe survey estimates firm AI adoption across business functions and shows that use remains narrow within most adopting firms; it does not identify causal productivity.
- National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile” (risk-management actions for testing, deployment decisions, incidents, and monitoring), 2024. nvlpubs.nist.govThe NIST generative AI profile adapts the AI RMF to risks specific to generative systems, which this chapter translates into operating controls.
Comments
Log in to comment