September 26, 2026
AI gets another archive of old texts: Oxford gave OpenAI 125,000 images
On Sep 26, The Guardian reported that OpenAI used digitized texts from Oxford's Bodleian Library to expand its training dataset. By June 2025, Oxford had provided the company with 125,000 images from Global Dissertations.

In March 2025, Oxford described the project as digitizing public-domain materials for researcher discovery and access. Internal documents show the data was also used to train OpenAI models.
ChatGPT runs on OpenAI models, and the Oxford project reveals another way the company is expanding their training data.
OpenAI is funding the digitization as part of a five-year partnership with Oxford.
