The AI race continues unabated. It’s interesting to note that Facebook appears to be ramping up their web scraping. Is this their plan B, in response to the EU not being happy about their plans to scrape their own users’ content without their consent? 1 They’ve been banging on my web server’s door harder than anyone else for the last few days. I’ve had them blocked for some time, so I can only assume their recent escalation of probing is their desperation for LLM training data.
I wonder when legislators will figure out that big tech is basically aiming tractor beams at the rest of the Internet. I can’t help but think of the giant eletromagnet on Lockdown’s “Knight Ship” from the “Transformers: Age of Extinction” movie, indiscriminately vacuuming up every magnetic object in its reach.

I have more than a decade’s worth of data for packets that have entered or exited my home network. I recently archived all but the last 2 years or so in cold storage. But I can say with confidence that the traffic to my web site has completely transformed in the last few years. It used to be mostly what appeared to be ordinary people (coming from residential broadband address space), most referred by a search engine or a message board. Today, the legitimate traffic looks the same as old, but it is dwarfed (by more than 4 decimal orders of magnitude) by big tech scraping and ne’er-do-wells using cloud infrastructure for nefarious purposes.
The Internet is quickly becoming another casualty of corporate greed. My self-written firewall automation is currently blocking 1,098,452,334 IPv4 addresses from port 80 and 443. That’s 27.54% of the publicly routable IPv4 unicast address space according to my bar napkin arithmetic. Yikes. Would you swim in a river where more than 1 in 4 of the molecules is toxic?
