Digitise Now, Train Later: Internal Documents Show OpenAI Is Using Oxford's Bodleian Library in Its Training Set
FOI-obtained documents reveal that texts digitised from Oxford's Bodleian Library have entered OpenAI's model-training set, sparking staff concerns over transparency and reputational risk.
What began in March 2025 as a five-year partnership to “digitise” Oxford’s world-famous Bodleian Library has turned out to be something more: according to internal university documents obtained via a freedom of information request and reported by the Guardian on September 26, 2026, the historical texts OpenAI has been scanning from Oxford’s collections are being used to “populate the OpenAI training set.”
The disclosure lands at a moment when frontier AI labs are scouring every credible source of human-written text — academic libraries, secondhand bookshops, university archives — because the open web has become saturated with AI-generated material that degrades model quality when fed back into training. Oxford is the only UK member of OpenAI’s NextGenAI consortium, which has struck similar agreements with the Boston Public Library, Caltech, MIT, and the University of Michigan. But it is the first to have its internal deliberations exposed in such detail, and the picture they paint is one of a venerable institution quietly negotiating the terms under which eight centuries of accumulated knowledge becomes fuel for someone else’s commercial models.
What the documents show
The core revelation is simple but consequential: when Oxford announced the partnership eighteen months ago, the public framing emphasised digitisation and access — using OpenAI software to make the Bodleian’s texts more widely available to students and researchers. What the announcement did not state, and what the FOI-obtained meeting minutes now confirm, is that the digitised material flows into OpenAI’s model training pipeline.
The scale is already substantial. By June 2025 — barely three months into the collaboration — 125,000 images scanned from historical dissertations had been shared with OpenAI, including PhD theses from European and American universities written in the 19th and 20th centuries. Beyond the dissertations, the scanned material includes a rare collection of 10,000 16th-century “broadside ballads” — song lyrics and musical notation once hawked on Tudor street corners — a category of text that exists almost nowhere else in digital form and is precisely the kind of fresh, human-written, pre-internet data that labs now prize most.
And it may not stop there. The contract opens the prospect of mass digitisation across the Bodleian’s full collection of 23 million items. Internal minutes also record discussions about creating an “Ask the Bod” chatbot built on the library’s holdings — a product concept that would put a conversational interface on top of five centuries of collected knowledge.
Staff pushed back internally
The meeting minutes obtained through the FOI request document concerns raised by university staff, including members of the Bodleian governance committee, on two main fronts.
The first is reputational risk. Partnering with the company behind ChatGPT — an organisation facing ongoing scrutiny over everything from agent misbehaviour incidents to regulatory disputes — ties a 700-year-old institution’s name to one side of a contested industry debate. Academics who spent 2025 and 2026 arguing about the ethics of AI training data now find their own employer contributing to the corpus.
The second is environmental. Staff questioned how a deal involving an energy-intensive technology squares with the university’s own sustainability commitments. Data centres and model training have become significant and growing energy consumers, and the minutes show that Oxford’s internal deliberations treated this as a genuine tension rather than a footnote.
Oxford’s defence: modest, out-of-copyright, non-exclusive
A University of Oxford spokesperson pushed back firmly against the suggestion that anything had been concealed. The amount of text being digitised is “modest in scale,” they said, and covers only out-of-copyright material. The Bodleian retains the rights to the scans and will begin publishing them openly online within months, as it does with outputs from other digitisation partnerships.
The spokesperson also rejected the claim that the machine-learning element had been hidden from the public and students: digitisation was the university’s primary interest, but staff had been open that the project would also contribute training data. OpenAI’s use of the material is non-exclusive, meaning other AI developers could in principle license the same scans. The university frames the project as a net win for access — the material will become available to a far wider audience than could ever physically visit the reading rooms on Broad Street.
An OpenAI spokesperson, for their part, said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future,” adding that with more than a billion people using the technology in everyday life, “it’s important it reflects different cultures, histories and perspectives.”
The wider data gold rush
The Oxford story is a window into a broader and increasingly desperate search for high-quality training data. Scraped websites are now heavily contaminated with AI-generated content, which creates a feedback loop — models trained on model output get worse, not better. Labs have responded by turning to physical and historical collections that predate the web entirely.
The tactics vary in tone. OpenAI’s approach — partnership agreements with libraries, with the books left intact on their shelves — is the genteel version. Anthropic, its close rival, has reportedly spent tens of millions of dollars acquiring secondhand books and slicing off their spines so the contents can be scanned before the books are pulped (the company says it does not buy and destroy rare or antiquarian volumes). Investigative outlet 404 Media went as far as placing a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where books were dismantled and scanned. Secondhand bookshop owners have speculated about why buyers snap up obscure titles — a guide to agricultural implements in 18th-century Africa, biographies of 1950s racing drivers — that have no plausible resale value except as unique, undigitised text.
Notably, the Bodleian’s collections remain physically intact under the OpenAI deal. Whatever else can be said about the arrangement, no Oxford volumes are being destroyed for the sake of a training corpus.
Why it matters
Three takeaways stand out.
Transparency is becoming the battleground. The Bodleian deal was announced as a digitisation and access project; the training-set usage surfaced only through an FOI request eighteen months later. As more universities weigh similar agreements — and NextGenAI’s expansion suggests many will — the difference between “digitisation partnership” and “training data partnership” is exactly the kind of distinction the public, and university staff, will demand be made explicit up front.
Pre-internet text is the new oil. Broadside ballads, 19th-century dissertations, Irish state papers, Dorothy Hodgkin’s penicillin notebooks — these are on the internal list of materials discussed for digitisation. They represent human knowledge that is nearly impossible to obtain any other way, and every exclusive scan of them is a competitive asset in the data arms race.
Institutions have leverage — if they use it. Oxford kept the rights to its scans, secured open publication within months, and made OpenAI’s use non-exclusive. That is a materially better position than the secondhand book market, where sellers have no idea who the buyer is or what happens to the books. The Bodleian case suggests the template other libraries may follow — or improve upon, by disclosing the training-data term in the original announcement rather than leaving it to be discovered.
The quiet question underneath all of it: once the world’s great libraries have been scanned, what’s left? The data gold rush that began with the web is now working its way backward through history — and Oxford, the oldest university in the English-speaking world, has just shown us what that looks like from the inside.