
Scrape.do's team explains why it avoids browsers on most requests, maintains sessions for difficult targets, tests approaches as sites change, and worries about pages that look successful but contain incomplete data.
From the outside, a scraping API looks simple. You send a URL. You get a page back.
Behind that one response, the Scrape.do team describes a stack of decisions, and several of them run against what most developers would assume.
The company reports that more than 99% of its requests work without a browser. Yet its engineers still open individual difficult targets and tune configurations for them by hand. And the failure they find hardest to catch is not the obvious block. It is the response that comes back successfully with part of the data missing.
That raises a question worth asking even if you only scrape a handful of sites: what does an operation handling 10 billion+ pages a month know about getting reliable data that still applies at your scale?
As Scrape.do's founder Batuhan Özyön put it during our recent Reddit AMA:
At a small scale, scraping is mostly a technical problem. At our scale, it becomes a continuous systems and operations problem.
Batuhan has worked in web scraping and reverse engineering for many years and has been building Scrape.do since 2020. He was joined in the thread by Lead Software Engineer Mert B., R&D engineers Raif Tekin and Muhammet Derviş Aygan, and Selman Gokce, who leads marketing and SERP product marketing. The traffic volumes, percentages and internal systems described below are the team's own reported figures, not independent measurements.
A disclosure: ScrapeOps routes a significant amount of traffic through Scrape.do, and we know the team well. That is exactly why we wanted to get past the feature list and into the operation behind it.
In our tenth r/WebScrapingInsider AMA, we asked the Scrape.do team when browsers are actually necessary, how sessions and site-specific configurations decide the hard cases, and what it takes to keep data flowing after the first request succeeds.
Here are the nine biggest insights from the discussion.