California requires training data to be disclosed
A requirement took effect to publish a description of training data: sources, quantity, types, legal status, presence of personal information and of synthetic data.
Why it matters
For the first time a developer must say what a model was trained on; until then that was a trade secret by default.
The law covers any system available to residents of the state and requires a description for every release since 2022. The list of required items is detailed: dataset owners, the number of data points in ranges, whether the material is under copyright, whether the data was bought or licensed, whether personal data is present, what modifications were applied. The exemptions are narrow: security systems, aviation, federal national defence. The practical consequence is not in the penalties but in this: litigants in training-data cases no longer have to prove what a corpus contained, because it has been published.