It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.
It is the “selling shit back to us” specifically, not the “scraping,” that’s the unethical part. If the AI companies were doing the same scraping (and destructive rare book scanning, for that matter), but were using the data to populate archive.org, would it still be a problem? I would argue “no.”
The “scraping” part becomes unethical when the scraping is so aggressive that it takes down the website (or severely impacts its ability to serve actual clients).
Archive.org scrapes the web all the time, but it doesn’t do it so aggressively that it becomes an issue for the websites they’re scraping. The same cannot be said for AI scrapers.
Scraping more than necessary is so stupid that I just sort of dismissed it as a straight-up mistake that will eventually be corrected. I was arguing based on general principle, not specific current practice.
Obviously, yes, the AI companies should fix their (probably vibe-coded) scrapers so they stop misbehaving; that should’ve gone without saying.
I think it would be a problem because the scrapers are hammering all types of websites from small forums to reddit with tens of thousands of unique ip addresses at a time. Websites that have neither the money, hardware, or protection had to figure out solutions really quick or suffer what is essentially a constant ddos attack. This is the reality of the web now, it’s just an incredibly hostile place.