AI Infra
0%
Part XI · Chapter 79

Data Rights and Compliance Economics

AuthorChangkun Ou
Reading time~13 min

Technically reachable data is not automatically rights-ready. Nor is a missing license an automatic legal conclusion. Rights-readiness is a dated decision about a particular data snapshot, proposed use, actor, jurisdiction, model or product release, and review date. It is not a property of the bytes.

That distinction changes the economics. Data can be copied and reused without being consumed, but access, contractual grants, privacy permissions, and some legal rights can still be exclusive or conditional. The valuable asset is therefore not merely a large corpus. It is a corpus tied to a defensible decision, working controls, and evidence that can survive a release review. This chapter explains an operating model for that decision; it is not legal advice. The legal classification itself belongs with qualified counsel and the jurisdictional workflow in Chapter 61.

Freeze the use before reviewing the data

"Can we use this dataset?" is underspecified. The answer can differ across pretraining, fine-tuning, evaluation, retrieval, output display, redistribution, logging and feedback, or generation of synthetic and derived assets. It can also change with the product, customer class, territory, duration, and party doing the processing.

Begin with a concrete use statement: this asset version, collected by this actor, will be processed in this way, for this product and release, in these jurisdictions, during this period. Then make admission an explicit gate:

Admit(a,u,j,r,t)=P(a)B(a,u,j,t)O(a,u,j,t)C(a,u,j,r,t)E(a,u,j,r,t).\operatorname{Admit}(a,u,j,r,t) = P(a) \land B(a,u,j,t) \land O(a,u,j,t) \land C(a,u,j,r,t) \land E(a,u,j,r,t).

Here, aa is the asset snapshot, uu the proposed use, jj the jurisdiction, rr the model or product release, and tt the review time. PP means provenance and lineage are complete enough for review. BB means an approved legal basis, permission, or exception covers the stated facts. OO means relevant obligations, rights reservations, and contract terms are recorded. CC means the required technical control is implemented. EE means current release evidence links the decision to the asset and build. This is an organizational release gate, not a universal legal test. A failed term means deny, quarantine, or send for review; it does not invite an engineer to invent a legal conclusion.

Separate evidence from the decision

A useful ledger keeps five layers distinct:

  1. Source facts: source URI, collector, date, content hash, asset version, transformations, and derivation. Provenance says where an asset came from; it does not prove permission.
  2. Legal decision: ownership or other legal basis, license, contract, relevant statutory exception, privacy basis, jurisdiction, reviewer, and source version.
  3. Organizational policy: customer commitments, excluded source classes, approval thresholds, and named risk acceptance. Policy can be stricter than law without becoming law.
  4. Technical control: allowlists, collection rules, filters, access control, attribution, retention, deletion, and release blocking.
  5. Release evidence: archived source material, signed terms, review record, control test, corpus manifest, model linkage, approver, and expiry date.

Collapsing these layers produces false certainty. A license label can be wrong. A contract can grant only some uses. A technical block can enforce policy without describing the legal basis. Conversely, a legal decision is not operational until a tested control carries it into the corpus, index, model workflow, and release.

data_rights_ledger facts Source facts identity · lineage · version use Proposed use actor · product · region facts->use decision Legal + policy decision basis · duties · approver use->decision controls Technical controls collect · retain · delete decision->controls evidence Release evidence manifest · tests · expiry controls->evidence
Figure 79.1. A rights decision moves from observed facts to release evidence. Any material change sends the asset back to review.

Signals are not interchangeable

A license, terms of service, a copyright rights reservation, privacy consent, and an access control answer different questions. The Robots Exclusion Protocol is a crawler-preference protocol. RFC 9309 says it is not access authorization and is not a substitute for security (Koster et al. 2022). Record a robots.txt rule and its collection time as an observed signal, not as ownership proof, privacy consent, or a complete legal decision.

Jurisdiction also matters. Article 4 of the EU Copyright in the Digital Single Market Directive makes an appropriately expressed, machine-readable rights reservation relevant to its text-and-data-mining exception (European Parliament and Council of the European Union 2019). That does not turn every crawler rule into a worldwide prohibition. Terms of service, copyright, license scope, privacy consent or another privacy basis, and access control remain separate and are not interchangeable.

Two audits show why the evidence layer matters. The Data Provenance Initiative examined more than 1,800 text datasets---1,858 in all---within 44 popular alignment-fine-tuning collections. Across three hosting platforms, more than 70 percent of license fields were unspecified in parts of the sample, while platform labels agreed with manual review only 35--54 percent of the time---an error rate of more than 50 percent in parts of the audit (Longpre et al. 2023). These are findings about selected dataset metadata, not legal conclusions about every dataset.

Consent in Crisis studied 14,000 web domains represented in C4, RefinedWeb, and Dolma. By April 2024, terms restricted roughly 45 percent of C4 tokens; fully restrictive crawler rules covered a smaller but rapidly growing share. The paper reports what would change if respected or enforced, and notes that an absence of restrictions does not imply consent (Longpre et al. 2024). Its observed signals are not legal conclusions, and a site operator may not own every work on the site.

Price a rights-ready option, not a token pile

Rights work has costs, but it can also create options: a market can open, a buyer's procurement gate can clear, or a model can ship sooner. Compare a candidate dataset DD with the best feasible baseline over one accounting horizon HH:

NBH(D)=ΔVH(D)Clicense(D)Cclear(D)Ccontrol(D)Cevidence(D)Cmonitor(D)E[ΔLH(D)].\begin{aligned} NB_H(D) ={}& \Delta V_H(D) \\ &- C_{\mathrm{license}}(D) \\ &- C_{\mathrm{clear}}(D) \\ &- C_{\mathrm{control}}(D) \\ &- C_{\mathrm{evidence}}(D) \\ &- C_{\mathrm{monitor}}(D) \\ &- \mathbb{E}[\Delta L_H(D)]. \end{aligned}

Here, NBHNB_H is incremental net benefit. ΔVH\Delta V_H is the change in product value, model performance, delivery speed, or markets and customer classes enabled. ClicenseC_{\mathrm{license}} is acquisition and royalty cost; CclearC_{\mathrm{clear}} is legal and operational clearance; CcontrolC_{\mathrm{control}} implements obligations; CevidenceC_{\mathrm{evidence}} produces reviewable records; and CmonitorC_{\mathrm{monitor}} covers renewal and change detection. E[ΔLH]\mathbb{E}[\Delta L_H] is the change in expected legal, contractual, privacy, remediation, and reputational loss relative to the baseline.

Put every term in the same currency and accounting horizon, amortize shared costs, and state the baseline. Avoid double counting a market-access benefit as both revenue and avoided loss. Keep non-negotiable privacy, safety, and legal limits as guardrails rather than assigning a convenient price to them. A license fee is not the dataset's intrinsic value or a per-token price: it reflects a negotiated bundle of content, access method, permitted uses, service, exclusivity, term, warranties, and bargaining power.

Reddit's 2024 prospectus illustrates that bundle. It disclosed several January 2024 data-licensing arrangements with a $203.0 million aggregate transaction price, two- to three-year terms, and at least $66.4 million of expected 2024 revenue. Substantially all contract value came from one partner, which the filing did not name (Reddit, Inc. 2024). The disclosure establishes a market transaction, not a unit price, a Google-specific price, or proof that every corpus has comparable value.

Court outcomes do not supply a global price

In June 2025, two Northern District of California judges issued record-specific district-court rulings, not a nationwide rule. Bartz v. Anthropic treated copies used to train the models as fair use on that record, but separated pirate-sourced books kept in a permanent general-purpose library and left that use for trial (Alsup 2025). Kadrey v. Meta entered summary judgment for Meta on the named authors' reproduction-and-training claim, yet emphasized their failure to produce meaningful evidence of market harm and analyzed the shadow-library downloads in light of the asserted training purpose (Chhabria 2025). The analyses differ; their facts, plaintiffs, markets, and procedural posture do not establish a universal rule for other models, media, outputs, or records. Neither order supports the shortcut that training is fair whenever material is lawfully acquired.

Bartz settled before the remaining issue was tried. On July 20, 2026, the court finally approved a $1.5 billion non-reversionary class settlement covering 482,460 listed works (Martínez-Olguín 2026). The agreement resolves specified past claims through August 25, 2025; it does not license future conduct or release output claims. A settlement is not precedent, an appellate holding, an admission of liability, or an adjudicated price for training.

Compliance can create market access

For providers of general-purpose AI (GPAI) models in the EU, Article 53 of the AI Act requires technical documentation, information for downstream providers, a copyright policy and compliance process, and a public training-content summary in the Commission's template (European Parliament and Council of the European Union 2024). Chapter V obligations began applying on 2 August 2025 to models placed on the market from that date; Commission enforcement, including the relevant Article 101 fines, began on 2 August 2026. Providers of models placed on the market before 2 August 2025 have until 2 August 2027 to comply.

The GPAI Code of Practice is a voluntary compliance tool, not the law itself. Its transparency and copyright chapters support Article 53 duties; its safety and security chapter addresses providers of GPAI models with systemic risk (European Commission 2025). Regulation (EU) 2026/1744 moved specified high-risk system deadlines, but did not postpone the GPAI dates (European Parliament and Council of the European Union 2026). The details, exceptions, roles, and transition calendar are developed in Chapter 61.

Documentation does not guarantee compliance. Economically, however, current evidence and working controls can make a product eligible for customers, regions, and regulated procurement paths that would otherwise remain closed. Measure that incremental access; do not relabel paperwork itself as value.

Open source answers a different question

The Open Source Initiative's Open Source AI Definition 1.0 requires freedoms to use, study, modify, and share an AI system. Its preferred form for modification includes code, parameters, and Data Information detailed enough for a skilled person to build a substantially equivalent system (Open Source Initiative 2024). That includes descriptions of training data, provenance, selection, processing, and available locations. It does not necessarily require publication of the training dataset itself when underlying data cannot be redistributed.

Weights alone therefore do not meet that definition, and an open source classification does not settle whether acquisition and processing complied with copyright, contract, privacy, or sectoral duties. Openness, reproducibility, and rights-readiness overlap, but one is not evidence of the others.

Operate the ledger through releases

Do not store a global usable = true. Store an allow, deny, or review decision against a versioned tuple of asset, use, product, region, customer class, and time. At minimum, the ledger and its tests should cover:

  • identity: source URI, snapshot hash, collection agent and date, transformations, derived datasets, indexes, and models;
  • decision: legal basis and reviewer, grantor authority, license or contract, rights reservation, privacy basis, and organizational policy;
  • scope: pretraining, fine-tuning, evaluation, retrieval, display, redistribution, logging, feedback, and synthetic or derived use by product and recipient;
  • obligations: territory, duration, attribution, payment, share-alike, redistribution, warranty, indemnity, audit, retention, deletion, and termination;
  • enforcement: collection and access control, filters, attribution, expiry blocks, deletion propagation, evidence tests, approver, and next review.

Re-run admission when source terms change, a rights reservation appears, a contract expires, deletion or withdrawal arrives, a new use, model, release, or new market is proposed, or a vendor change alters the processing chain. Deletion also needs a scoped meaning: deleting a source record, removing it from future corpora and retrieval indexes, and remediating an already trained model are different operations; see Chapter 59.

What's contested

More restrictive access can improve payment, provenance, and creator control while raising entry barriers and concentrating valuable corpora among firms able to fund large deals. Broader exceptions can support research, competition, and new products while leaving creators and downstream buyers to bear more uncertainty. The disagreement is not only over whether training creates value. It is over which rights apply, who may grant them, who pays for clearance and enforcement, and how those choices distribute market power.

Constraint arrow

The ledger begins with the collection and quality controls in Chapter 6 and the lineage, privacy, retention, and deletion machinery in Chapter 59. Its legal decisions come from Chapter 61. Upstream, it determines which products and regions are available to the market analysis in Chapter 77. Rights evidence is therefore part of deployable capability, not metadata added after a model is built.

Further reading

  • Longpre et al., “The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI” (Audit of licensing and attribution metadata in 1,858 datasets), 2023. arXiv:2310.16787
    The Data Provenance Initiative audits more than 1,800 text datasets and finds widespread license omissions, license errors, and a divide between commercially open and closed data.
  • Longpre et al., “Consent in Crisis: The Rapid Decline of the AI Data Commons” (Longitudinal audit of observed restriction signals across 14,000 web domains), 2024. arXiv:2407.14933
    Tracks changing robots and terms-of-service signals across domains represented in major web corpora while separating those signals from legal consent.
  • Koster et al., “Robots Exclusion Protocol” (Protocol specification stating that robots rules are not access authorization), 2022. rfc-editor.org
    Standardizes robots.txt as a crawler-preference protocol and explicitly distinguishes it from access authorization.
  • Alsup, “Bartz v. Anthropic PBC, Order on Fair Use, No. 3:24-cv-05417” (District-court order separating training copies from pirate-sourced books retained in a general-purpose library), 2025. govinfo.gov
    Finds the challenged model-training use fair on its record while leaving the separate permanent-library use of pirate-sourced books for trial.
  • Chhabria, “Kadrey v. Meta Platforms, Inc., Order on Summary Judgment, No. 3:23-cv-03417” (District-court fair-use order emphasizing the named plaintiffs' deficient market-harm record), 2025. law.justia.com
    Grants Meta summary judgment on the named authors' reproduction-and-training theory while stressing the absence of meaningful market-harm evidence.
  • European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence” (Official Journal text of the EU AI Act), 2024. eur-lex.europa.eu
    Sets the binding EU framework, including Article 53 duties for providers of general-purpose AI models and the transition dates in Article 113.
  • European Commission, “The General-Purpose AI Code of Practice” (Voluntary compliance tool for transparency, copyright, and systemic-risk safety and security duties), 2025. digital-strategy.ec.europa.eu
    Provides a voluntary route for demonstrating specified GPAI transparency, copyright, safety, and security obligations.
  • Open Source Initiative, “The Open Source AI Definition 1.0” (Open-source AI freedoms and the preferred form for modification), 2024. opensource.org
    The OSI definition treats Open Source AI as requiring use, study, modification, and sharing freedoms, with data information, code, and parameters available in the preferred form for modification.
  • European Commission, “Guidelines for Providers of General-Purpose AI Models” (Nonbinding Commission guidance on GPAI model scope and provider obligations), 2025. digital-strategy.ec.europa.eu
    Explains the Commission's intended interpretation of GPAI model and provider concepts under the AI Act.

Comments

Log in to comment