Training "AI" On Public Data Is Totally Fine And Not Stealing.

@[email protected] · 4 months ago

Training "AI" On Public Data Is Totally Fine And Not Stealing.

@[email protected] · 4 months ago

Incidentally, I read this a while ago, because I was training a classifier on mostly Creative Commons licensed works: https://creativecommons.org/2023/08/18/understanding-cc-licenses-and-generative-ai/

… we believe there are strong arguments that, in most cases, using copyrighted works to train generative AI models would be fair use in the United States, and such training can be protected by the text and data mining exception in the EU. However, whether these limitations apply may depend on the particular use case.

@[email protected] · edit-2 4 months ago

Maybe there should be a distinction if an individual does is for educational and research and a corporation does it for commercial use. As a user it’s fun and usefull to generate whatever mix of text or images I want from a model that was trained on everything, but a user doesn’t see the exploitation made by the corporation that handed him the tool