Why Your Scraper Keeps Breaking, and What Reliable Data Delivery Looks Like
You paid a freelancer to build a scraper. It filled your spreadsheet for about three weeks, then one morning the rows stopped appearing and nobody could tell you why.
If that sounds familiar, you are not unlucky and you did not hire an idiot. The question of why web scrapers break has a short list of boring, predictable answers, and almost all of them trace back to the same root cause: someone wrote a script that worked once and called it a system. Those are not the same thing, and the gap between them is exactly where your data went.
Let me walk through what actually breaks, why the cheap version dies on a schedule, and what a data feed looks like when it is built by someone who expects it to run for years instead of weeks.
The short list of why web scrapers break
A scraper is just a program that visits web pages and copies specific information[1] off them into something you can use, like a spreadsheet or a database. If you want the longer version, I wrote a plain-English guide to what web scraping is. The part that matters here: a scraper depends entirely on things it does not control. That is the whole problem in one sentence.
Here are the usual killers.
The website changed. A scraper finds data by looking for landmarks in the page's code, like "the price is in the box labeled product-price." When the site's team renames that box, moves it, or redesigns the page, your scraper looks for a landmark that no longer exists and comes back with nothing. The site owner has no idea you exist and no reason to warn you. They ship a Tuesday redesign and your feed goes dark.
The site started fighting back. Larger sites run anti-bot systems, which are tools designed to tell a real human visitor apart from an automated one and block the automated ones. If your scraper requests pages too fast, in too regular a rhythm, or in a way that looks nothing like a person clicking around, it gets served a puzzle, a fake page, or a flat "access denied."
Your address got blocked. Every request your scraper makes carries an IP address[2], which is like a return address on an envelope. Send a few thousand envelopes from the same address in an hour and the site starts refusing anything from it. A throwaway script usually runs from a single address, so it is the easiest thing in the world to shut off.
The data moved behind a wall. Sometimes the information you want only appears after a login, a search, or a button click, and a simple scraper never sees it. Getting at that reliably is its own discipline, and a script that was never built for it just hands you an empty file with no explanation of what it missed.
Something upstream hiccuped. The site was briefly down. Your server ran out of memory. The network dropped one request in fifty. On its own each of these is trivial. To a script with no plan for failure, any one of them is fatal.
Notice the theme. None of these are exotic. They are Tuesday. A system that runs in the real world has to assume all of them will happen, often on the same day, and keep delivering anyway.
A script is not a system
This is the distinction that decides whether you get data next month.
A script is a set of instructions that runs top to bottom and stops. Someone points it at a site, it works on their laptop, they hand you a clean file, everyone is happy. The trouble is that a script has no memory, no second chances, and no way to tell anyone when it fails. It assumes every page loads, every landmark is where it was yesterday, and nothing ever says no. The first time reality disagrees, the script either crashes or, worse, quietly returns garbage and keeps going.
A system assumes the opposite. It expects pages to fail, layouts to move, and addresses to get blocked, and it is built to notice, absorb the hit, and recover without a human staring at it. That difference is not a nice-to-have. It is the entire product. When people compare the cost of web scraping and a service quote looks expensive next to a $200 freelancer, this is what the gap pays for. The freelancer sells you the script. The engineered feed sells you the part that survives contact with the real internet.
A demo proves the code can succeed once. Production is the craft of failing gracefully ten thousand times a day and still handing you the row.
Why "it worked in the demo" is a trap
The demo is where cheap scrapers look identical to good ones, which is exactly why it fools people.
On demo day the site has not changed since yesterday, the volume is tiny, you have not tripped any rate limits, and one clean address is doing all the work. Every condition that eventually kills the scraper is absent. Of course it works. A scraper working in a demo tells you as much about its durability as a car starting in the driveway tells you about a cross-country drive.
Reliability does not show up at small scale. It shows up at volume, over time, under pushback. A run of a hundred pages once proves nothing. A run of a hundred thousand pages a day, every day, for six months, through three site redesigns and a new anti-bot vendor, is the only test that counts. Judge a scraper by its worst week in month four, not its best hour on day one.
What reliable data delivery actually requires
When a data feed keeps running, it is not because the code is smarter. It is because a few unglamorous systems are doing their jobs in the background.
Retries that are patient, not stubborn
When a page fails, the system should wait and try again, and it should wait longer after each failure instead of hammering the site (which just gets you blocked faster). It should also know the difference between "try again in a minute" and "this page is genuinely gone, stop wasting effort." Blind retries are almost as bad as no retries. The intelligence is in knowing which failures are worth a second attempt.
Monitoring and alerting
The system has to watch its own output and know what normal looks like[3]. If yesterday brought in 40,000 records and today brought in 12, something broke, even if nothing technically crashed. Good monitoring catches the quiet failures, the ones where the scraper runs happily and returns half the data or subtly wrong data. And when it catches one, it tells a human immediately, before you find out from a customer or a bad decision made on stale numbers.
Self-healing
The best systems route around problems on their own. Address blocked? Rotate to a fresh one[4] and keep going. A page format the scraper does not recognize? Flag that record, set it aside for review, and keep processing the other 99,000 instead of falling over. The goal is that most breakages get absorbed automatically, and only the genuinely new ones ever reach a person.
I lean on all three of these hard on the Injuria platform I built, a legal-intelligence system that processes more than 500,000 pages a day. At that volume something is always failing somewhere. The platform stays useful not because nothing breaks, but because a break in one corner never takes down the whole feed, and because it tells me what needs attention instead of making me go looking.
What a real ongoing data feed looks like
If you are buying data as a service rather than a one-time file, here is the standard worth holding a provider to.
- Freshness you can count on. You know how current the data is and when the next refresh lands, and that promise holds on a bad week, not just a good one.
- Uptime, not heroics. The feed keeps delivering through site changes and blocks because recovery is built in, not because someone happened to be awake to babysit it.
- A named human who fixes it. When something does break, and eventually something will, there is a specific person whose job is to notice and repair it, usually before you do. That accountability is most of what you are paying for, and it is the first thing I tell people to check in my notes on hiring a web scraping partner.
The mistake that puts people in the rebuild-every-quarter loop is treating a living data feed as a one-time purchase. Websites are not static, so nothing pointed at them can be either. The cost of a scraper is not the day it gets built. It is every day after that it has to keep working.
If your data keeps drying up and you are tired of being the one who discovers it went dark, I can help. Send me the sites you need watched and how fresh the data has to be, and I will tell you straight whether it needs a real system or just a cleanup, and what keeping it alive actually takes. You can grab a time at calendly.com/itsmattgeorge.