cyberivy
OpenAIUniversity of OxfordBodleian LibraryAI Training DataLibrariesNextGenAIAI Transparency

Oxford let OpenAI train models on Bodleian texts

September 26, 2026

Ein Smartphone zeigt das OpenAI-Logo vor einer von DALL-E erzeugten Landschaft mit Bäumen und Wasser.

Internal documents show that digitized Bodleian Library texts entered OpenAI's training set. Public announcements had initially emphasized digitization and research access.

What this is about

The University of Oxford did more than digitize historical texts from its Bodleian Library with technology from OpenAI. According to internal university documents obtained by the Guardian through a freedom-of-information request, the material was also used to populate OpenAI's training set. The newspaper published its investigation on September 26, 2026.

When OpenAI and Oxford announced their partnership through the NextGenAI consortium in March 2025, the public emphasis was different: rare texts would be transcribed and made searchable for researchers. The fact that the digitized material would also supply training data was much less prominent. That gap between digitization and model training is what makes the story important.

What the collaboration actually does

OpenAI supplied technology to scan and transcribe public-domain historical collections. According to the Guardian, 125,000 images of historical dissertations had been shared by June 2025. The material also included 10,000 sixteenth-century broadside ballads. Internal minutes describe the digitized content as contributing to OpenAI's training set.

Oxford says the material is modest compared with a total collection of roughly 23 million items, is non-exclusive and is out of copyright. The Bodleian Library retains rights to the scans and says it intends to publish them openly as well. OpenAI argues that historical sources can help modern models reflect a wider range of cultures, periods and perspectives.

Why it matters

High-quality, human-curated text is becoming more valuable to model developers as the open web contains an increasing amount of machine-generated material. Libraries, by contrast, hold reviewed collections, metadata and rare works that are still missing online. The case shows how the role of public knowledge institutions is changing: a digitization project can simultaneously become a source of raw material for a commercial model.

For students, researchers and the public, the issue is therefore not limited to copyright. Transparency, participation, access rights and the question of who benefits commercially from publicly or charitably held knowledge all matter. According to the Guardian, internal Oxford minutes also recorded concerns about reputational risk and energy use. There was little visible public debate about those issues when the partnership began.

Libraries will need to explain such contracts more precisely: which files a partner receives, the purposes for which they may be processed, and whether later model versions may benefit as well. A verifiable timetable for open access matters too, so the public does not merely supply the source material but receives a tangible benefit.

In plain language

Think of the library as a public kitchen with a rare recipe archive. A company helps photograph the handwritten recipes and make them searchable. If it also uses those photographs to train its own cooking machine, that is not automatically wrong. But visitors should be told clearly in advance that the digitization project also produces training material.

A practical example

Suppose a library scans 125,000 pages from old dissertations. Researchers gain full-text search and can find names or concepts in minutes rather than weeks. The same collection also enters the training of a language model. The model may become better at processing historical writing, but outsiders cannot determine how much those 125,000 pages affect a particular model or whether scanning errors are carried into it.

Scope and limits

  • The Guardian investigation establishes that the material entered the training set, but it does not show a measurable effect on any specific OpenAI model.
  • Oxford says only public-domain works are involved and that the use is non-exclusive. This is therefore not equivalent to unauthorized training on copyrighted books.
  • Internal meeting minutes document concerns within the university, but they do not necessarily represent every employee, student or library user.

It also remains unclear how precisely Oxford will document datasets, selection criteria, scanning errors and specific model uses. Without that information, the scientific and public benefits cannot be fully weighed against the risks.

SEO & GEO keywords

University of Oxford, Bodleian Library, OpenAI, training data, NextGenAI, historical texts, library digitization, public-domain works, AI transparency, academic data

💡 In plain English

Oxford digitized historical public-domain library texts and also supplied the results to OpenAI for model training. The texts are meant to become publicly accessible, but the training purpose was not prominent in the original announcement.

Key Takeaways

  • →Digitized Bodleian texts became part of OpenAI's training set, according to internal documents.
  • →By June 2025, the shared material included 125,000 images of historical dissertations.
  • →Oxford stresses that the works are public domain, the use is non-exclusive and the collection is small relative to its total holdings.
  • →The library retains rights to the scans and plans to publish the digitized material openly.
  • →The effect of the material on individual models remains unknown.

FAQ

Which texts were provided to OpenAI?

The disclosed material includes historical dissertations, sixteenth-century broadside ballads and other public-domain collections. The complete selection has not been documented publicly in detail.

Are the works copyrighted?

Oxford says only public-domain material is involved. The Bodleian Library also retains the rights to the scans it created.

Will the digitized material be publicly accessible?

The university says the material will begin to be published openly in the coming months.

Do we know which model was trained on it?

No. The documents refer to OpenAI's training set but do not identify a specific model or a measurable effect.

Sources & Context