things changed, and they want a piece of the cake now.
they get the chance to adapt to new stuff that's happening and profit from it, however the only ones left without a say are the users that generated that content.
I think it would be fair to say that previously accepted terms and conditions shouldn't ethically count for this, and user data previously generated should be shared for AI training only on an opt in basis.
> user data previously generated should be shared for AI training only on an opt in basis.
This seems obviously backwards and a boon to enormous corporations.
If you want to build a search engine, you have to index everything possible, or it won't be able to find that and then it's useless. Transacting with each individual person in the world would only be possible for megacorps -- it's already expensive enough to index everything if you don't have to do that. It's also perverse that someone could have a right to prevent someone else from providing true information about them. If you're standing there making a false claim, you shouldn't have a right to prevent someone else from proving you wrong just because the evidence is from your own past.
But search engines have been ML since even before LLMs, and LLMs are fundamentally the same. How can someone have a default right to deny the public access to facts?
Now, maybe there are some things you want to keep private, and then if you share them with someone you want to bind them to not sharing that information with others, like an NDA. But that's opt-out, not opt-in, and you can't really do that for things that are public.
Is it? It's not "like an NDA". You get the data erased (if it is agreed that your privacy is more important than the public interest) after the fact of the data being made public.
Which is why it's controversial, and inconsistent with free speech.
If you want to be forgotten, you change your name and move somewhere else. That should be a right, because it puts it into the hands of the person who wants it to happen. You should be able to e.g. change your social security number, without the government telling anyone what your old one was. Now you're "forgotten" and get to start over.
How do you have a right to selectively suppress information about your past? That's just a boon to liars and scammers because they can do it continuously at no cost to themselves.
And even then, it's still something you have to explicitly ask for, not something that happens by default, notwithstanding that you can do it retroactively.
Right, getting the data erased is the opt-out part. Opt-in would be if they had to give permission to crawl in the first place, which is what folks are suggesting for AI training.
> user data previously generated should be shared for AI training only on an opt in basis.
Wouldn't all of the populate models and applications (chatgpt, copiolot, etc) have to be burned to the ground and start clean? My understanding was that pretty much every one of these were trained on "stolen" data, i.e. without prior agreement with the authors / content creators.
I think the law was clear before that any data on public internet which doesn't require signing the TOS or login is publicly scrapable. People have been doing it for decades for things like market sentiment analysis and even reselling the data for the same. Why would AI wave suddenly change it.
they get the chance to adapt to new stuff that's happening and profit from it, however the only ones left without a say are the users that generated that content.
I think it would be fair to say that previously accepted terms and conditions shouldn't ethically count for this, and user data previously generated should be shared for AI training only on an opt in basis.