Oxford sent OpenAI 125,000 Bodleian theses to train models
TL;DR
- By June 2025 the Bodleian had sent OpenAI 125,000 scans of 19th- and 20th-century PhD theses from European and US universities.
- Internal Oxford documents describe the digitised material as used to 'populate the OpenAI training set', a use the public March 2025 announcement did not disclose.
- Staff including members of the Bodleian governance committee raised reputational risk and questioned the fit with Oxford's environmental commitments.
The Bodleian Libraries sent OpenAI 125,000 scans of old PhD theses by June 2025, mostly 19th- and 20th-century work from European and US universities, and internal Oxford documents describe the material as being used to "populate the OpenAI training set" — a use the university's March 2025 announcement did not spell out, The Guardian reports.
That announcement framed the project as an access effort, saying OpenAI software would digitise Bodleian holdings and make them more widely available to students and researchers. Oxford is the only UK member of OpenAI's NextGenAI group, whose other members include Boston Public Library, Caltech, MIT and the University of Michigan. Three of the AI researchers we follow circulated the story the day it ran.
Meeting notes obtained through a freedom of information request show staff, including members of the Bodleian governance committee, raised the reputational risk of the OpenAI tie-up and questioned how a partnership with an energy-intensive technology squared with the university's environmental commitments.
An Oxford spokesperson said "the material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI's use of the material is not exclusive," adding that the library retains the rights and "will begin to publish these materials openly online in the next few months." An OpenAI spokesperson said: "With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives."
Shared on Bluesky by 3 AI experts
Originally reported by theguardian.com
Read the original article →Original headline: Oxford lets OpenAI train its AI models on Bodleian Library