this post was submitted on 10 Jul 2023
67 points (100.0% liked)

Technology

23 readers
2 users here now

This magazine is dedicated to discussions on the latest developments, trends, and innovations in the world of technology. Whether you are a tech enthusiast, a developer, or simply curious about the latest gadgets and software, this is the place for you. Here you can share your knowledge, ask questions, and engage in discussions on topics such as artificial intelligence, robotics, cloud computing, cybersecurity, and more. From the impact of technology on society to the ethical considerations of new technologies, this category covers a wide range of topics related to technology. Join the conversation and let's explore the ever-evolving world of technology together!

founded 2 years ago
 

In addition to the possible business threat, forcing OpenAI to identify its use of copyrighted data would expose the company to potential lawsuits. Generative AI systems like ChatGPT and DALL-E are trained using large amounts of data scraped from the web, much of it copyright protected. When companies disclose these data sources it leaves them open to legal challenges. OpenAI rival Stability AI, for example, is currently being sued by stock image maker Getty Images for using its copyrighted data to train its AI image generator.

Aaaaaand there it is. They don’t want to admit how much copyrighted materials they’ve been using.

you are viewing a single comment's thread
view the rest of the comments
[–] Ferk@kbin.social 15 points 1 year ago (2 children)

Note that what the EU is requesting is for OpenAI to disclose information, nobody says (yet?) that they can't use copyrighted material, what they are asking is for OpenAI to be transparent with sharing the training method, and what material is being used.

The problem seems to be that OpenAI doesn't want to be "Open" anymore.

In March, Open AI co-founder Ilya Sutskever told The Verge that the company had been wrong to disclose so much in the past, and that keeping information like training methods and data sources secret was necessary to stop its work being copied by rivals.

Of couse, disclosing openly what materials are being used for training might leave them open for lawsuits, but whether or not it's legal to use copyrighted material for training is something that is still in the air, so it's a risk either way, whether they disclose it or not.

[–] 00@kbin.social 9 points 1 year ago

and that keeping information like training methods and data sources secret was necessary to stop its work being copied by rivals.

Cant have others copying stuff that you have painstakingly copied yourself.

[–] nicetriangle@kbin.social 4 points 1 year ago

They seem really intent on having their cake and eating it too.

a) we're not violating the letter or spirit of copyright laws

b) disclosing our data could open us up to a ton of IP lawsuits

hmm