The Irony of 'You Wouldn't Download a Car' Making a Comeback in AI Debates

FatCat@lemmy.world · 13 days ago

The Irony of 'You Wouldn't Download a Car' Making a Comeback in AI Debates

Hackworth@lemmy.world · 13 days ago

Equating LLMs with compression doesn’t make sense. Model sizes are larger than their training sets. if it requires “hacking” to extract text of sufficient length to break copyright, and the platform is doing everything they can to prevent it, that just makes them like every platform. I can download © material from YouTube (or wherever) all day long.

castlebravo404@lemmynsfw.com · 13 days ago

They’re absolutely not doing everything they can. Everything they can would be to not use the works. They’re doing as much as they’re willing to do. If it wasn’t for the threat of lawsuits they wouldn’t even be doing that much.

mm_maybe@sh.itjust.works · 13 days ago

Model sizes are larger than their training sets

Excuse me, what? You think Huggingface is hosting 100’s of checkpoints each of which are multiples of their training data, which is on the order of terabytes or petabytes in disk space? I don’t know if I agree with the compression argument, myself, but for other reasons–your retort is objectively false.

beebarfbadger@lemmy.world · 13 days ago

The issue isn’t that you can coax AI into giving away unaltered copyrighted books out of their trunk, the issue is that if you were to open the hood, you’d see that the entire engine is made of unaltered copyrighted books.

All those “anti hacking” measures are just there to obfuscate the fact that that the unaltered works are being in use and recallable at all times.