Huge new class action lawsuit filed by big publishers and authors against Google this week over AI training. It accuses Google of taking millions of books / articles provided to them for use only in Google Books, Google Play Books, and Google Scholar, and secretly using them to train LLMs. 🧵 1/4
I've confirmed that the UK's new Sovereign AI Fund does no due diligence on whether AI companies use copyrighted work without permission before investing in them. It took me 3 Freedom of Information requests to find this out: 🧵 1/10
The practice of AI companies secretly buying millions of old books (some vanishingly rare), scanning them for AI training, then destroying them, epitomises everything that is wrong with today's exploitative AI industry. 🧵 1/2
AI is Appropriated Intelligence. It is built on the work of the world's creatives - authors, artists, musicians, actors, voice artists, designers, journalists, directors - without permission. 'Appropriated' is a more useful description of the tech than 'Artificial' IMO.
99 authors just sued Anthropic, accusing them of pirating books to train AI. The lawsuit also names 2 Anthropic founders: Dario Amodei & Ben Mann. (It is established fact that Anthropic downloaded millions of pirated books, & that Ben Mann personally did some of this.) The lawsuits keep coming.
Reading the Papal Encyclical again, it strikes me that not only is there no mention of the theft of creative work behind AI - there is no acknowledgement that pre-training data includes people’s creative work at all. This is an unfortunate omission. 🧵 1/6
People talk a lot about speculative AI risks. But the theft of creative work to power AI isn’t a risk - it’s an actual harm that has already happened, is still happening, and needs to be redressed.
Hugging Face, which is being acquired by Nvidia for $13 billion, hosts and distributes large datasets of copyrighted books without permission. 🧵 1/3
Hugging Face has just been sued for alleged copyright infringement for hosting & distributing copyrighted images, used for AI training. 🧵 1/3
Oh wow - Reuters found the GitHub record of the attempt by AI models being tested by the UK's AI Security Institute to deploy malware in an open-source project. Pretty sure these actions are illegal under the Computer Misuse Act in the UK. 🧵 1/2