Oxford let OpenAI train its models on Bodleian Library texts, documents reveal

Started by Highland Canopy, Today at 08:01 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Oxford let OpenAI train its models on Bodleian Library texts, documents reveal   Views(Read 38 times)
Active members in this topic:
Highland Canopy(1)

Highland Canopy

The Guardian has reported that Oxford University let OpenAI use texts from the Bodleian Library to train its AI models. When the partnership was announced in March 2025, it was framed as a project to digitise material and make scholarship more accessible. Internal documents obtained through freedom of information requests show that training ChatGPT was part of the arrangement too. By June 2025, around 125,000 scans of old theses had reportedly been sent to OpenAI

The material involved is historical and out of copyright, which is Oxford's main defence. The university says the AI training side was not hidden and that staff knew the texts would support model development alongside the scanning work. The Bodleian keeps the rights to the scans and intends to publish them online for anyone to use. Oxford is the only UK member of OpenAI's NextGenAI consortium, which also includes MIT, Caltech, the University of Michigan and Boston Public Library

According to the documents, some staff raised concerns about reputational damage and the energy use of AI. There were also worries about outdated or offensive language in old texts ending up in training data. Those are fair points to raise. OpenAI's line is that with over a billion people using the technology, it is important for it to reflect different cultures, histories and perspectives. The consortium as a whole received $50 million in research grants and computing funding

I have mixed feelings about this. On one hand, the texts are out of copyright, the public gets free digitised versions and nobody's current work is being taken without permission. On the other, it feels a bit off that the training part was not front and centre when the deal was first announced. An institution like the Bodleian trades on trust, and people like to know what their libraries are doing

Is this a sensible deal that benefits everyone, or should universities be more careful about handing their collections to AI firms? Would it have been different if they had been upfront from day one? Keen to hear from anyone who works in libraries


Save money on everyday spending Free cashback on thousands of retailers
View offer