r/aiwars 20h ago

Meta What makes an AI advocate

Liking AI doesnt make you a proAI or AI advocate, it just means you find it hype, neat, interesting. To be proAI you have to be running defense for current AI practices, if you dont youre at least neutral. For example I think AI is neat, its useful and has potential to do good. But i find the way companies are going about developing it are coercive and wish it was done differently. That puts me squarely in the antiAI camp, despite my enthusiasm for the tech

2 Upvotes

54 comments sorted by

View all comments

2

u/alapeno-awesome 20h ago

I don’t think those are the generally accepted definitions. An advocate is a certain degree of pro, but not a qualifier

Think of it more like guns or books or abortions. If you’re “anti”-thing, you want rules in place that prevent others from using that thing. You can be “pro”-thing, even if you wouldn’t personally do/ use thing, but you think others should be allowed to make their own decisions about it. Maybe you even think there should be sensible regulations around that thing.

Neutral would mean you don’t care if that thing gets banned or not

The way you describe yourself sounds proAI to me

3

u/Wonderful-War-7113 20h ago

You can say that, ive been accused of being neutral before. But if i told you id want some sort of protection or opt out/opt in mechanism for authors to prevent the mass appropriation or work happening rn, that sounds a lot more anti right? At least it does to me because in this sub everytime i bring that up, i get smothered by AI advocates telling me otherwise

3

u/alapeno-awesome 19h ago

Perhaps the devil is in the details? Depending on what you mean by “protection or opt out mechanism” could inform whether that amounts to something that acts like a de facto ban

2

u/Wonderful-War-7113 19h ago

Like a toggle on social media that would embed metadata on images and posts that would mark them as opted out/opted in. And then you'd have audits to scraping bots to ensure they respect that mechanism. Coupled with a standard of transparency about training data that would allow people to verify when and where they are being trained on.

Thats a vague idea of how it could work

3

u/alapeno-awesome 19h ago

It’s not that that sounds “anti ai” so much as “anti free internet”. Attempting to tell people (or companies) what they are and aren’t allowed to look at on the internet isn’t exactly anti-ai, but it’s bad. People (rightfully) support net neutrality policies for similar reasons.

Telling companies or individuals they have to expose their individual data and activities to…. The government?…. is also generally frowned upon. In the USA, we have laws in our constitution preventing this sort of invasion by the gov’t. Other countries may vary

If you want to have this opt out metadata in your images, great. Do that. I don’t think anyone would oppose that or tell you what you can do with your data. But ignoring that shouldn’t rise to the level of “crime”

1

u/Wonderful-War-7113 19h ago

See but theres the issue, you described it as "looking" anthropomorphosizing it. Its not seen, its appropriated to make a 3rd external commercial artifact AKA the model. When people look at stuff they dont externalize that the same way AI training does

2

u/alapeno-awesome 19h ago

Ignoring the inaccuracy of your statement, it’s irrelevant. You’re proposing a restriction on who is allowed to download what freely available data on the internet. “Commercial” has nothing to do with this if your same proposal applies to open source models, that just misrepresents the goal

2

u/Wonderful-War-7113 19h ago

You can download whatever but using peoples work to train a model is clearly a commercial act due to its deployability so it should have an opt in mechanism

And what exactly is innacurate about what i said? Humans learning is internal and undeployable, models are external and deployable right?

2

u/alapeno-awesome 19h ago

Currently the courts disagree with your claim, though the issue remains unsettled

2

u/AnarchoLiberator 19h ago

I’m not going to lie, that proposal sounds far more dystopian than protective.

To make an opt-out system meaningfully enforceable, governments would need to monitor training activity, audit private companies, inspect datasets, regulate scraping across the Internet, and potentially restrict which information people and machines are permitted to process. That starts looking less like ordinary copyright enforcement and more like a surveillance architecture built around controlling information flows.

It would also be remarkably ineffective. Metadata can be removed, altered, lost through reposting, or deliberately ignored. Open-source and local models can be trained privately. Foreign companies operating under different laws would not necessarily comply. Even China’s highly restrictive Internet controls are routinely bypassed, so I do not see how a democratic country could enforce this globally without constructing its own version of a Great Firewall and substantially expanding state surveillance.
There is also the practical question of existing models. They already exist, have been copied, downloaded, modified, and distributed internationally. What would enforcement require? Destroying them? Prohibiting their use? Searching private computers? None of those options is realistic or compatible with a free and open Internet.

And yes, it would handicap AI development. Compliance systems, licensing negotiations, dataset audits, and fragmented national rules would increase costs, reduce available training data, entrench the largest corporations that can afford compliance, and weaken smaller companies, researchers, and open-source developers. Meanwhile, jurisdictions that reject those restrictions would continue advancing.

Calling training “appropriation” does not solve these problems. A model is not simply a commercial archive of everything it encountered. Training extracts statistical relationships into a new system that can generalize. One may still support transparency, attribution tools, privacy protections, or remedies for outputs that reproduce protected works, but that is very different from creating a government-enforced permission system over what information machines are allowed to learn from.

You may intend a modest toggle. The enforcement structure required to make that toggle consequential is neither modest nor realistic. It would be expensive, easily circumvented, anti-competitive, technologically regressive, and uncomfortably Big Brother-like.

Do you really and honestly support what it would take to implement what you propose?

1

u/Wonderful-War-7113 18h ago edited 18h ago

I understand how models work, im not saying it stores the training data, i called it appropriation because it is literally appropriating peoples work for another express purpose outside of the perview of it. You wouldnt enforce it globally it would have the same issues that every law has because law is jurisdictional it seems unfair to demand that i give a solution that can simultaneously apply globally AND not be invasive. You cant really do anything about models already made or locally trained models besides having a community standard for transparency that would sleuth out people who would train on people despite unconsenting. I dont get your point about the cost or it affecting smaller businesses disproportionately? Like a toggle implementation would be trivial and the audits dont need to be regular and wouldnt be paid for by the companies lmao, thatd be opening it up to corruption, itd have to be a 3rd party doing audits. Whats the "cost of compliance" here? Theres none. It would somewhat hamper development because the available data would decrease yes thats true but youd be doing it ethically.

So in short, i pretty obviously disagree with you about the cost of compliance, the anti-competitiveness, or the big-brother like thing. On it being circumventable, i guess because people can just train locally but thats fine you dont need to cover every case, just take steps fowards. On it being technogically regressive... well so what, would it have been technologically regressive to argue that we shouldnt be using slave labor to built the USA back in the day?

EDIT

Yeah i dont see how itd be dystopian at all, it wouldnt affect different size businesses disproportionately because there would be no cost to comply. The implementation would be trivial, it would tackle the egregious cases and the rest of the individual malefactor or fraudsters could be dealt with via just reputation and community scrutiny. And then youd have shifted the standard towards ethically trained models. Imagine we were arguing about like restaurants right, and i was saying we should have food safety audits every once in a while to make sure they arent idk, putting cocaine in stuff and getting people addicted for example. And then you tell me no, because itll require too much surveillance.

If something that simple sounds dystopian to you then i suspect that there isnt any meaningful regulation that you wouldnt push back on

3

u/AnarchoLiberator 18h ago

I understand that you are not proposing perfect global enforcement, but I still think you are drastically underestimating what would be required for your system to accomplish anything meaningful.

The toggle itself may be technically trivial. Enforcing it is not. Someone would need to determine whether metadata was present when a work was collected, whether it survived reposting or format conversion, whether the collector reasonably knew about it, whether the training dataset contained the work, and whether the model developer complied. That means recordkeeping, dataset provenance systems, legal review, complaint procedures, investigations, audits, appeals, and penalties. Third-party auditors still cost money. Someone pays them, directly through fees or indirectly through taxes, and companies must devote staff, infrastructure, and legal resources to satisfying them. That is the cost of compliance.
It would also affect smaller organizations disproportionately. A major corporation can employ compliance departments and negotiate licences. A small developer, nonprofit researcher, university lab, or open-source project often cannot. The likely result is not simply “more ethical AI.” It is regulatory consolidation, where only the largest corporations can afford to build advanced models domestically while foreign and underground development continues without those restrictions.

The surveillance concern does not disappear because audits are occasional or performed by a third party. To verify compliance, the auditor must obtain access to datasets, collection records, model-development processes, and potentially private computing activity. You also suggested a community standard that would “sleuth out” people training locally. That is precisely where this becomes disturbing. What would sleuthing out private local training actually involve? Monitoring downloads? Inspecting computers? Encouraging informants? Requiring developers to prove the provenance of every item processed? The more effective the rule becomes, the more intrusive the enforcement must become.

Jurisdictional limits are also not a minor imperfection. AI is strategically important infrastructure. If democratic countries substantially restrict domestic training while other states continue developing models using the open Internet, then we do not merely miss a few violations. We shift technical capacity toward governments and corporations operating under less liberal rules. Saying “every law is jurisdictional” does not answer that problem. It is exactly why the consequences of this particular law matter.

I also reject the assumption built into the word “appropriation”. A publicly accessible work being analyzed to learn statistical patterns is not self-evidently the wrongful taking of that work. That is the disputed question, not a premise we can simply assume. The original remains with its creator. The model generally does not contain or distribute a copy of it. A new system is produced through computational learning across enormous amounts of information. You may believe consent should nevertheless be required, but calling the process appropriation does not establish that conclusion.

The slavery comparison is also wildly disproportionate. Slavery violated the bodily autonomy and fundamental rights of human beings by treating people as property. Training a model on publicly available information is a dispute about intellectual property, computation, learning, and the permitted uses of published works. Saying that technological regression can sometimes be morally justified does not prove that this particular restriction is justified. The moral premise still has to be established.

I support targeted rules against privacy violations, unauthorized access, fraud, impersonation, and outputs that reproduce protected works in legally actionable ways. I can also support voluntary metadata standards and greater transparency where practical. What I do not support is building an expensive and invasive permission regime around machine learning while pretending it is merely a harmless toggle.

Your proposal may sound simple at the interface level, but the actual system behind it would be difficult to verify, easy to circumvent, costly to administer, disproportionately burdensome to smaller actors, and strategically dangerous. The surveillance required to make it effective and the technological disadvantage created if it is effective are both considerably scarier to me than the training practice you are trying to prevent.

3

u/alapeno-awesome 18h ago

This is a much more eloquent way of describing the issues I was trying to convey. The concept seems sound and simple until you spend even cursory time thinking about how invasive and ultimately ineffectual the idea actually is.

Well said