Novelists count what is inside Books2
On 28 June 2023 Paul Tremblay and Mona Awad filed a class action against OpenAI, case 3:23-cv-03223. OpenAI has never disclosed what the Books1 and Books2 sets hold, so the complaint sizes them from figures in the GPT-3 paper: Books1 about nine times BookCorpus and Books2 about forty-two times, which works out at about 63,000 and about 294,000 titles.
Why it matters
This is the first suit by novelists over a language model's training data, and it showed a way of reasoning that would become ordinary: when the contents of a corpus are closed, you derive its size from the maker's own published numbers, and then ask where that many books could have come from.
The works named are Tremblay's The Cabin at the End of the World and Awad's 13 Ways of Looking at a Fat Girl and Bunny. The claims cover copyright, DMCA section 1202 and unjust enrichment. Nine days later, on 7 July, the same lawyers filed a second suit against OpenAI for Sarah Silverman, Christopher Golden and Richard Kadrey, case 3:23-cv-03416; the cases were later consolidated. Paragraph 35 of that second complaint does what the first does not: it names the shadow libraries Library Genesis, Z-Library, Sci-Hub and Bibliotik and says Books2 holds books copied from them. The phrasing is on information and belief, an inference from size rather than an established fact, and the report presented it as recorded in the filings. What the record does not claim. The figures 63,000 and 294,000 are the plaintiffs' estimate, not OpenAI's data. Nothing is said about how the case ended.