this post was submitted on 11 Jul 2023
156 points (98.8% liked)

Piracy: ꜱᴀɪʟ ᴛʜᴇ ʜɪɢʜ ꜱᴇᴀꜱ

54500 readers
345 users here now

⚓ Dedicated to the discussion of digital piracy, including ethical problems and legal advancements.

Rules • Full Version

1. Posts must be related to the discussion of digital piracy

2. Don't request invites, trade, sell, or self-promote

3. Don't request or link to specific pirated titles, including DMs

4. Don't submit low-quality posts, be entitled, or harass others



Loot, Pillage, & Plunder

📜 c/Piracy Wiki (Community Edition):


💰 Please help cover server costs.

Ko-Fi Liberapay
Ko-fi Liberapay

founded 1 year ago
MODERATORS
 

cross-posted from: https://lemmy.world/post/1330512

Below are direct quotes from the filings.

OpenAI

As noted in Paragraph 32, supra, the OpenAI Books2 dataset can be estimated to contain about 294,000 titles. The only “internet-based books corpora” that have ever offered that much material are notorious “shadow library” websites like Library Genesis (aka LibGen), Z-Library (aka B-4ok), Sci-Hub, and Bibliotik. The books aggregated by these websites have also been available in bulk via torrent systems. These flagrantly illegal shadow libraries have long been of interest to the AI-training community: for instance, an AI training dataset published in December 2020 by EleutherAI called “Books3” includes a recreation of the Bibliotik collection and contains nearly 200,000 books. On information and belief, the OpenAI Books2 dataset includes books copied from these “shadow libraries,” because those are the most sources of trainable books most similar in nature and size to OpenAI’s description of Books2.

Meta

Bibliotik is one of a number of notorious “shadow library” websites that also includes Library Genesis (aka LibGen), Z-Library (aka B-ok), and Sci-Hub. The books and other materials aggregated by these websites have also been available in bulk via torrent systems. These shadow libraries have long been of interest to the AI-training community because of the large quantity of copyrighted material they host. For that reason, these shadow libraries are also flagrantly illegal.

This article from Ars Tecnica covers a few more details. Filings are viewable at the law firm's site here.

you are viewing a single comment's thread
view the rest of the comments
[–] SinJab0n@mujico.org 22 points 1 year ago (1 children)

I'm ok with PEOPLE reading books in any way for self improvement.

But, when a FUCKING COMPANY starts screwing with shit like this, thats when they crossed the line.

[–] Vendetta9076@sh.itjust.works 12 points 1 year ago (1 children)

Sure but you understand that publishers dont give a fuck about any of that. They find any way to shut these things down they can. Not to mention the things on Sci-Hub and Libgen should be free public knowledge to anyone or anything that wants it. Its full of tax funded research papers and textbooks. That information should belong to everyone and everything. Thats not a crossed line. Thats consistency.

[–] SinJab0n@mujico.org 3 points 1 year ago (1 children)

I agree with u, it should be free to every PERSON who wants it.

As i said before thats the fundamental difference between individuals and a company stealing.

[–] Vendetta9076@sh.itjust.works 1 points 1 year ago

We dont agree. Its not stealing and companies should have access to the same free information.