Skip to main content

Why Diagnosis Beats More Stealth: Lessons From 7 Years of Reverse Engineering Anti-Bot Systems

· 25 min read

Why Diagnosis Beats More Stealth: Lessons From 7 Years of Reverse Engineering Anti-Bot Systems

From the AMA

Experienced scraping engineers get better results by diagnosing what actually failed, keeping sessions consistent and choosing complexity carefully, not by stacking on more stealth.

Your scraper gets blocked. So you swap the proxies.

Still blocked. You add a stealth plugin, then a new browser, then a CAPTCHA solver.

Each change feels like progress. But each one is a guess, and some of those guesses make things worse. A scraper that rotates its fingerprint every few requests can get blocked faster than one that changes nothing.

The engineers who get consistent results start somewhere else. Before they change anything, they work out what actually failed, and ask one question: does this change address the actual cause?

As freelance anti-bot engineer Ibrahim El Khalil Mlata put it during our recent Reddit AMA:

Ibrahim is a web scraping and anti-bot engineer based in Algeria who has spent the last seven years reverse engineering anti-bot defenses across retail pricing, legaltech, logistics, hospitality and AI-training-data pipelines. Among the work he has done: running 1,000+ spiders across 300 retailers, reviving a dead 160-spider fleet, matching 100K+ hotel reviews a month, and taking mobile apps apart with Frida when the website is locked down. You can find his work on GitHub.

In our eleventh r/WebScrapingInsider AMA, we asked Ibrahim how he diagnoses blocks, when reverse engineering is worth the effort, and what keeps a scraping system useful once it is in production.

Here are the nine biggest insights from the discussion.


More Stealth Can Make Your Session Less Believable

1. More Stealth Can Make Your Session Less Believable​

One community member described a familiar experiment. They had read that rotating TLS/JA3 fingerprints helps avoid detection, so they switched fingerprints every few requests on a small scraper.

It got blocked faster than when they left it on one fingerprint and didn't touch anything.

Ibrahim's explanation started with a line worth remembering:

Too much stealth is a signal in itself.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The reason is simple once you think about what a real visitor looks like. A browser's TLS fingerprint comes from the browser and operating system. It doesn't change halfway through a visit.

Real users dont change their TLS fingerprint every few requests, it stays the same for the whole session because its tied to their browser and OS. So when the same IP or cookies show up with a new fingerprint every few requests, thats something no real browser does and you stand out more than if you did nothing

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

So the randomization didn't hide the scraper. It produced an identity that no longer resembled one browser on one device. The IP and cookies said "same visitor". The fingerprint said "different client". That contradiction is the signal.

When we asked what mattered more than he initially expected, Ibrahim widened the lesson from TLS to the whole session:

What mattered more than I expected is consistency. Keeping everything aligned for the whole session, TLS, headers, user agent, browser profile, and IP all telling the same story. A clean consistent setup with a decent IP beats a perfect proxy with a messy fingerprint

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

That last sentence matters for budgets. A better proxy improves one signal. It can't fix a user agent that disagrees with the TLS fingerprint, or headers that don't match the browser profile.

This doesn't mean every fingerprint change is harmful, or that every defense scores sessions the same way. It is a specific failure pattern: variation added without asking whether it makes the session more coherent.


Identify What the Target Checks Before Choosing the Fix

2. Identify What the Target Checks Before Choosing the Fix​

A block tells you that something failed. It doesn't tell you what.

Several community members asked about mouse curves, typing jitter and behavioral simulation. Ibrahim's view was that on many targets he encounters, IP reputation and device, identity and browser fingerprints matter more, and that most of the time the run is lost before behavior is even checked.

But he was careful not to turn that into a universal rule. The type of defense changes where effort pays off:

What helps is detecting what kind of security system youre facing first. A regular anti-bot cares mostly about IP, TLS and fingerprints, while anti-fraud systems give a lot more weight to biometrics and human behavior. Once you know which one youre dealing with you know where to put your effort, and on anti-fraud targets behavior simulation is real risk reduction not wasted time

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Recognizing the vendor helps, but it isn't the whole picture. One participant was struggling with Akamai on airline sites, having tried Playwright, Patchright and Camoufox without consistent results. Ibrahim pointed out that the vendor name can hide a larger system:

Sometimes its not only Akamai. Big companies like airlines usually have their own security team, and Akamai is just one piece of their anti-bot stack along with captchas and other tools, all orchestrated together. Thats what makes it hard

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

So how does he start on a new target? With the cheapest possible test, then working down only as far as he needs to:

Start with website analysis, then first discovery, send a few curl requests and if it works mimic it in Python. If it doesnt, figure out why, is it TLS fingerprinting, browser fingerprinting, JS rendering etc and keep going down the funnel

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The part many teams skip is what comes after. Ibrahim records what he learns about each site as a profile:

Then you classify each website by different attributes like difficulty, anti-bot vendor, JS rendering, how often the site changes and so on. Next time you get a new website you walk it through the same pipeline, build a profile for it, and from that profile you get a pretty solid cost estimate

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Notice "how often the site changes" in that list. Whether you can access a target today is only half of the assessment. How often it will break is what drives long-term cost.


Solving the CAPTCHA Does Not Remove Every Detection Signal

3. Solving the CAPTCHA Does Not Remove Every Detection Signal​

When a browser scraper keeps hitting CAPTCHAs, adding a solver looks like the obvious fix.

One participant asked the harder version of the question: what if the challenge gets accepted, and the next request immediately triggers another one?

Ibrahim started from the challenge itself rather than the solver:

It depends on the captcha youre dealing with. Each one is sensitive to different things, some care about human behavior, some detect Chrome extensions, some lean heavily on browser fingerprinting

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Understanding what a particular challenge looks at tells you where to start. In his approach, that means improving things at the browser and human-behavior level first. Only then does he decide how to integrate a solver.

He described two options. You can use a solver's Chrome extension. Or you can build your own flow: get the solution from the solver service, then submit it yourself, with realistic mouse movement.

He usually prefers the second, and explained why:

I usually prefer the second option because solver extensions are detectable, and anti-bot engineers actively reverse engineer the solutions these solver services use, so relying on them out of the box can get you flagged even when the captcha itself is solved

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

That last clause is the insight. Whether the CAPTCHA answer is correct and whether the integration is detectable are two separate problems. A solver can give you the right answer while the way it was delivered, or the browser it was delivered from, still gives you away.

This is Ibrahim's preference based on his experience. It isn't a claim that every solver extension is detectable, or that a custom submission flow guarantees success. But it does explain why a high solve rate can coexist with a scraper that keeps getting challenged.


Reverse Engineering Pays When Savings Outlast the Maintenance Burden

4. Reverse Engineering Pays When Savings Outlast the Maintenance Burden​

Reverse engineering a bot manager has a certain appeal. Crack the sensor payload once, generate valid tokens without a browser, and scrape faster and cheaper from then on.

Ibrahim treats it differently:

For me deciding to reverse engineer a bot manager is more of a business decision than a technical one. The question isn't "can I reverse it?" but "is it worth reversing?"

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The main input to that decision is stability:

The main thing I look at is how stable the system is. If a bot manager stays the same for a long time then reverse engineering pays off, you do the hard work once and then you can generate valid tokens or payloads without a browser which is faster cheaper and scales way better

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

"Without a browser" doesn't always mean a plain HTTP client, though. Asked how he injects reversed sensor data, Ibrahim said it varies. Sometimes a simple JavaScript injection is enough. Sometimes you need a JavaScript engine to execute the script, paired with something like curl_cffi for the requests, with no full browser needed. Other times the whole solution needs to be architected around it.

The cost side of the ledger is real. Ibrahim said the initial work sometimes takes weeks, or even more, and that anti-bots are getting harder to reverse. Then there is the repair cost when the target changes:

If the system changes all the time its usually not worth it. Google SearchGuard is a good example, they push new versions constantly so anything you reverse breaks fast and you have to start over. You end up maintaining it forever instead of actually collecting data.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

In those cases, he said, a solid setup with good proxies, clean fingerprints and a real browser is the smarter move, even if it costs more per request.

Volume is the other half of the calculation. Runtime savings are per request, so they only add up if you send enough requests:

If the anti-bot is stable and the volume is high enough, reversing pays off.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

So it isn't only perfectly stable targets that justify reversing. It's whether the savings, at your volume, keep outrunning the cost of the initial work plus each repair. Ibrahim noted elsewhere in the thread that working bypasses usually last weeks or months, and that a vendor update is the most common reason they break.


A Different Interface May Be a Better Route to the Same Data

5. A Different Interface May Be a Better Route to the Same Data​

When a website is hard to scrape, the natural response is to add more browser complexity. Another stealth patch. Another automation framework.

In his answer about Akamai-protected airline sites, Ibrahim described a different first move:

Strategically, sometimes I jump to the mobile app instead of wasting time on the browser. If thats not possible, then investigate the anti-bot really well and understand what signals its collecting from your side.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The broader habit is reconnaissance. Before committing to one interface, check where else the data you need is exposed. A website, a mobile app and any other client may all serve the same data through different defenses.

"Different" isn't the same as "easier", though. Mobile apps come with their own protections and their own engineering burden. Ibrahim described a recent investigation where AI tooling helped him map exactly that:

A recent one was reversing a mobile app using JADX, Burp Suite MCP and Claude Code. It helped me investigate the app and quickly identify which protections it was using, root detection, Magisk detection, SSL pinning etc

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

If you want to explore this route, his mobile toolkit is a useful starting point:

For mobile apps, JADX to read the code, Frida to hook functions at runtime, Burp Suite for traffic, and Ghidra when the logic is in native libraries

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The AMA doesn't include a worked comparison showing that mobile was cheaper, faster or less protected for a particular target. The point is narrower and more useful: the browser is one route to the data, not the only one, and it's worth knowing your alternatives before you invest heavily in any of them.


Reduce Scraping Costs by Removing Unnecessary Browsers and Proxy Requirements

6. Reduce Scraping Costs by Removing Unnecessary Browsers and Proxy Requirements​

When scraping costs climb, the instinct is to shop around for a cheaper version of the same infrastructure. Cheaper proxies. A cheaper browser service. A cheaper API.

Asked where a small team overspending on difficult sites should look first, Ibrahim pointed at the setup itself:

I think the best place to start is understanding your scraping setup really well. That means doing an audit, collect all your performance data and analyze it.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

His examples of what to look for were concrete. Which scrapers run on a browser but could switch to plain HTTP? Which use residential proxies when datacenter ones would work just fine?

These are easy configurations to inherit. A scraper gets a browser because the target needed one two years ago. A pool gets set to residential because it was the safe default. Nobody revisits either decision, and both keep costing money.

Ibrahim's advice is to check with data rather than guess:

Collect as much data as you can, analyze it, and where youre not sure, run AB tests before changing anything. You can do this once and then repeat it once a year or so, kind of like a KYC but for your scraping project. Most of the savings usually come from there, not from switching to a cheaper tool

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

This connects with something he flagged as overstated advice:

I think people overspend on proxies. Sometimes we just put more money and effort into better proxies when the real problem is the browser fingerprint or the way were scraping.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

None of this means browsers or residential proxies are generally unnecessary. Some targets need both. The point is that many setups are never re-checked against what the target needs today, and that's where the unclaimed savings tend to sit.


AI Speeds the Build, Not the Moving Target

7. AI Speeds the Build, Not the Moving Target​

The "AI made one engineer 10x" narrative came up directly. One participant asked whether one engineer can now build and maintain substantially more scrapers.

Ibrahim's answer was measured:

Yes those claims are mostly hype. AI made me faster, but not 10x

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The speed-up is real. He said Claude Code, along with MCPs, skills and tools he has set up around it, made the biggest difference to his workflow in recent years: site analysis and discovery, debugging and building the scraper are all faster.

In reverse engineering specifically, he uses it for the grind, and he was clear about where it stops helping:

I use Claude Code mostly for the boring parts, cleaning and renaming obfuscated code, explaining functions, writing Frida scripts or Babel transforms faster, and identifying which protections an app uses. It fails on custom VMs and anything that needs real understanding of the whole flow, and it can be confidently wrong, so you still need to verify everything yourself

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

That verification step is why experience still matters. In Ibrahim's view, the engineers who benefit most are the ones who already know what to look for:

Yes, but mostly for an experienced engineer who knows how to use it and combine it with other tools, pentesting and RE tools like Frida, JADX etc. It saves me a lot of time on the investigation side. But LLMs dont have the knowledge that comes from real experience, so you still need the human intuition part, the gut feeling that tells you where to look

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Then there is maintenance. Ibrahim reports faster debugging with AI, which might seem to make maintenance easier too. His answer explains why it doesn't fully:

Building scrapers is a lot faster now, investigation, writing the code, debugging, parsing. But maintenance is still the hard part, sites change, anti-bots update, data breaks silently. That stays just as labor intensive and still needs a human who knows what hes doing

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

The two claims fit together. AI shortens each investigation. It doesn't stop the target from changing, the vendor from pushing an update or the data from quietly going wrong. The flow of maintenance work comes from outside your codebase, and each item still needs someone who can judge whether the fix is right.

These are Ibrahim's observations from his own work, not measured benchmarks. But they are a useful counterweight to the idea that AI turns scraping into a solved problem.


A Successful Request Can Still Produce Commercially Wrong Data

8. A Successful Request Can Still Produce Commercially Wrong Data​

A 200 response and a valid schema feel like success. They don't prove that the data describes the right product.

One participant running competitor pricing across six marketplaces described exactly this. For their team, the anti-bot side was the predictable part. What cost them money was a scraper coming back "successful" with a price matched to the wrong variant: 500ml versus 1L, bundle versus single unit, refurbished versus new. Nobody caught it until margins looked off two weeks later.

That's the participant's production example, not Ibrahim's. But his response drew a clear line between two kinds of validation. Asked whether to start with schema validation alone or add business rules from day one, he said both:

Id do both from day one, but keep the domain rules simple. Pydantic alone only tells you the data has the right shape, it wont catch a 1L price matched to a 500ml product, and thats exactly what costs you money

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Schema validation answers "is this a price?" Domain validation answers "is this the price of the thing we think it is?" A library like Pydantic can host custom rules too. The point is that basic field checks alone won't catch this class of error.

His starting set of rules was deliberately small:

So start with schema validation plus a few basic rules, like price change vs last week above a threshold, pack size or variant mismatch, and duplicates for the same EAN. Flag them instead of dropping them, review what gets flagged, and add more rules over time based on what you actually find

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

"Flag them instead of dropping them" is the detail to keep. Silently discarding suspicious records hides the problem. Flagging them puts the anomalies in front of someone who can tell a real price drop from a matching bug.

For a second line of defense, Ibrahim recommended checking against the source:

From my experience the best way to catch this is to compare the scraped data with whats actually live on the website. You can do it manually from time to time as a data quality check, just pick a random sample and verify it against the site.

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

He also noted that, unlike the anti-bot side, this part is fully in your control.


At Scale, Missing Debugging Evidence Becomes a Major Maintenance Cost

9. At Scale, Missing Debugging Evidence Becomes a Major Maintenance Cost​

With a handful of scrapers, a failure is an annoyance. With hundreds, failures are a constant stream, and the cost of each one depends on how quickly you can work out what happened.

We asked Ibrahim what started consuming his team's time once they were running fleets of hundreds of scrapers. His answer wasn't proxies or anti-bot updates:

What consumed the most time at scale was debugging. A lot of the time we didnt have enough information around the bug to know what to fix quickly

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

That is a different problem from "the bug was hard". The fix might have been easy. The cost was in not knowing which fix to make.

Another participant framed the same issue from the other side: a 403 on its own could be the IP, request rate, TLS fingerprint, headers, cookies, browser environment or prior session behavior. Without more context, you're guessing. And the usual guess, as Ibrahim noted in the answer that opens this article, is to replace the proxies, when in his experience it's often the fingerprint or the scraping behavior.

Isolating variables one by one only works if you can see the variables. What helped his team was more evidence:

What helped was collecting as many metrics as possible and having better logging. Now with AI and MCPs connected to Grafana, Loki or any other monitoring stack, debugging got a lot faster, the AI can go through the logs and metrics with you and point you to the problem much quicker

Ibrahim El Khalil Mlata · Freelance Anti-Bot Engineer

Our interpretation: the AI part of that answer depends on the first part. An assistant connected to your logs can only search what you recorded. If the logs don't capture what distinguishes one likely cause from another, the AI will be guessing too, just faster.


Conclusion: Diagnosis Keeps a Scraper Accessible, Correct and Maintainable​

Most of the conversation around anti-bot systems focuses on access: getting past the block. More stealth is the usual answer.

Ibrahim's answers point to diagnosis instead, and not only for access. Getting in depends on consistent sessions, an accurate read of what the target checks and complexity that is chosen on purpose rather than piled on. But a scraper that gets in and returns the wrong variant's price isn't working either. And a scraper that works today but can't be diagnosed when it breaks becomes an ongoing cost.

As Ibrahim summarized:

You can't stop targets from changing. You can control whether you understand why a fix should work, whether your data is checked for meaning, and whether you'll have the evidence to diagnose the next breakage.

Next time your scraper gets blocked, diagnose before you disguise: work out which layer failed, then choose the fix that addresses it.