I mean maybe they do have enough compute to run OCR with character-correction on every image format file on the internet. If we’re just speculating then we might as well speculate about that.
But like, seatbelts don’t prevent every death. We still wear them though, because the reduction of risk is still significant.
The fact is, catching an acronym appearing in images requires a much larger funnel than catching it in plaintext. And catching it with a white line through it on those images requires a larger funnel than without.
It reduces the attack surface. Maybe there’s still an attack surface, but if it’s much smaller then it’s still worth it.
You may aswell just censor it better, as said, Google’s OCR can easly detect it (i just tested lmao) it wouldn’t be hard at all for them to get the OCR text and scan for acronyms if they have enough resources
Yeah but the point is that trawling the entire internet and scanning every single image with OCR would take way more compute than just scanning particular images on demand which is what happens when you test it on google.
Scanning every jpeg on the internet with OCR would require far more compute than simply doing a regex for plaintext.
Also that white line can really fuck with most lightweight OCRs
Yeah but we are following their theory so i guess we just ignore that for now (I know it’s stupid and would take so much processing power)
I mean maybe they do have enough compute to run OCR with character-correction on every image format file on the internet. If we’re just speculating then we might as well speculate about that.
But like, seatbelts don’t prevent every death. We still wear them though, because the reduction of risk is still significant.
The fact is, catching an acronym appearing in images requires a much larger funnel than catching it in plaintext. And catching it with a white line through it on those images requires a larger funnel than without.
It reduces the attack surface. Maybe there’s still an attack surface, but if it’s much smaller then it’s still worth it.
You may aswell just censor it better, as said, Google’s OCR can easly detect it (i just tested lmao) it wouldn’t be hard at all for them to get the OCR text and scan for acronyms if they have enough resources
Yeah but the point is that trawling the entire internet and scanning every single image with OCR would take way more compute than just scanning particular images on demand which is what happens when you test it on google.