The Record — AI Data Law and Provenance | The Gray Systems
THEGRAY.SYSTEMS
← The Gray Systems

The Record

Laws, judgments, and settlements that determine how AI training data must be sourced, documented, and licensed. Entries are added only when a matter concludes. Nothing here is pending.

Every matter below turns on the same question, and it is not whether a model was trained. It is what the trainer can produce as evidence of where the corpus came from. In each case that evidence had to exist before anyone knew it would be needed. A training corpus cannot be documented in retrospect; the record is either created at collection or it does not exist. Each entry ends with the specific record it demands.

EU AI Act — transparency and general-purpose obligations

2 August 2026 · European Union

Article 50 transparency obligations apply and enforcement over general-purpose AI models begins. Providers publish a training-data summary on the Commission's template and honour copyright opt-outs expressed through machine-readable signals. Penalties reach €15 million or 3% of global turnover. Parliament moved most high-risk obligations to 2027 and 2028 in June 2026; these provisions were not moved.

An opt-out signal can only be honoured at the moment of collection, and can only be proven afterward by a log written at that moment.

The record it demands: per-source collection timestamp with the state of that source's machine-readable reservation as it stood at collection.

Bartz v. Anthropic

June 2025 ruling, settlement approved July 2026 · Northern District of California

Training on lawfully acquired books was held fair use. Downloading and retaining pirated copies was held not to be. The case settled for $1.5 billion, approximately $3,000 per work across roughly 482,000 works.

The court separated the training step from the acquisition step and found liability in only one of them. The training was free; the acquisition cost $1.5 billion.

The record it demands: an acquisition receipt per source, linking each work in the corpus to the transaction or licence that obtained it.

California AB 2013 — Generative AI Training Data Transparency Act

In force 1 January 2026 · California

Developers serving Californians since January 2022 must publish a summary of training datasets across twelve categories, including dataset sources and owners, intellectual property status with licensing details, whether personal information is present, and whether AI-generated synthetic data was used. There is no trade secret exemption. xAI has challenged the statute.

Synthetic data is named in the statute, which means a synthetic corpus now carries a public disclosure obligation rather than a private one.

The record it demands: declared lineage for generated material — what model produced it, under what licence that model was released, and what the inputs were.

hiQ v. LinkedIn and Meta v. Bright Data

2022 and January 2024 · Ninth Circuit and Northern District of California

The Ninth Circuit held that automated collection of publicly available data is not unauthorised access under the Computer Fraud and Abuse Act, and that a platform cannot place parts of its public site off limits to particular parties. The Northern District of California applied the same reasoning in Meta v. Bright Data, dismissing the CFAA claim; the contract claim survived only for the period Bright Data held an account. Separately, hiQ was held liable for engaging contractors to create false accounts to reach logged-in data.

Authentication is the line. Data visible without an account is collectible, and the moment a credential is used to reach it the question stops being about access and becomes about contract.

The record it demands: authentication state at collection — whether any credential, account, or session was used to reach each source.

Thomson Reuters v. Ross Intelligence

11 February 2025 · District of Delaware · on appeal to the Third Circuit, argued 11 June 2026

Westlaw headnotes and the Key Number System were held original and protected, and copying them to train a competing legal research tool was held not to be fair use. Market substitution was decisive, and the court weighed harm to a potential licensing market for training data.

The court treated a market for licensed training data as a thing that exists and can be damaged, which makes licensing the expected route rather than the cautious one.

The record it demands: for each source, the licence obtained, or documented evidence that no licensing market existed at the time of collection.

Kadrey v. Meta Platforms

June 2025 · Northern District of California

Training was held fair use regardless of whether the underlying material was legitimately obtained, reasoning differently from Bartz, decided days earlier in the same district. The court identified market dilution as a stronger theory of harm than the one argued. Claims concerning torrenting remain live.

Two judges in one district reached the same result on training and disagreed about acquisition, which means a supplier who can document acquisition does not need either line of reasoning to survive.

The record it demands: acquisition documentation held independently of any doctrinal position, sufficient under whichever reasoning prevails on appeal.

GEMA v. OpenAI

2025 · Munich Regional Court

The court held that the European text-and-data-mining exception does not extend to output a model has memorised.

A mining exception covers analysis, not reproduction, so an obligation taken on at collection continues into what the trained model is capable of emitting.

The record it demands: a retained manifest of what entered the corpus, sufficient to test any output against it.

Getty Images v. Stability AI

November 2025 · United Kingdom High Court

The court rejected the argument that model weights are themselves infringing copies. Getty succeeded on a narrow trademark claim concerning watermarks visible in model outputs, and was ordered to pay 69.4% of Stability's legal costs.

The weights proved nothing and the watermark proved everything, because the watermark was the only thing that traced an output back to a source.

The record it demands: a durable identifier carried from source item into the corpus, so any output can be traced back to what produced it.

Thaler v. Perlmutter

2 March 2026 · United States Supreme Court

The Supreme Court declined review, leaving in place the Copyright Office's refusal to register a work generated without human authorship.

Generated output carries no copyright of its own, so its entire defensible value sits in the documentation of how it was made.

The record it demands: the human contribution log — what a person did to the material, and when.

The Gray Systems builds training datasets to specification — synthetic and collected — with documented chain of custody and flexible licensing. Defensible by design.

Start Scoping Details

© TheGray.Systems . All Rights Reserved.

Legal
scroll