Tag: scrape protection

My experience of manual, no-code scrape of a bot-protected site

Post author By admin
Post date April 5, 2024
No Comments on My experience of manual, no-code scrape of a bot-protected site

Recently we discovered a highly protected site — govets.com. Since the number of target brand items of the site was not big (under 3K), I decided to get target data using the handy tools for a fast manual scrape.

Tags JSON, scrape protection, web scraping

Development

Amazon scrape tip

Recently we’ve met requirements to scrape Amazon data in big quantities. So, first of all I’ve tested the data aggregator for being bot-proof or anti-bot protection. For that I used the Discord server Scraping Enthusiasts, namely Anti-bot channel.

Since Amazon is a hige data aggregator we recommend readers to get acquainted with the post Tips & Tricks for Scraping Business Directories.

Tags anti-scrape, scrape detection, scrape protection

Challenge Development

Discord Bot to detect on-site anti-scrape & scrape-proof tools

Post author By admin
Post date March 22, 2023
No Comments on Discord Bot to detect on-site anti-scrape & scrape-proof tools

Today, I’ll share of a Dicord server 1 and server 2 that accomodate a bot able to detect multiple modern scrape-protection and scrape-detection means. The server’s channels with the bot are #antibot-test and #antibot-scan respectively

Tags anti-scrape, CloudFlare, scrape detection, scrape protection

Challenge

Bot protected websites

Tags anti-scrape, CloudFlare, scrape protection

Development

Headless Chrome detection and anti-detection

Post author By admin
Post date January 29, 2021
No Comments on Headless Chrome detection and anti-detection

In the post we summarize how to detect the headless Chrome browser and how to bypass the detection. The headless browser testing should be a very important part of todays web 2.0. If we look at some of the site’s JS, we find them to checking on many fields of a browser. They are similar to those collected by fingerprintjs2.

So in this post we consider most of them and show both how to detect the headless browser by those attributes and how to bypass that detection by spoofing them.

See the test results of disguising the browser automation for both Selenium and Puppeteer extra.

Tags anti-scrape, headless, Javascript, scrape detection, scrape protection

Development

How to find out that website is Distil protected?

Post author By admin
Post date November 6, 2020
No Comments on How to find out that website is Distil protected?

Given: a webpage to scrape.
If you inspect the DOM tree of that page you will find that quite a few tags are having the keyword dist. As an example:

<link rel="shortcut icon" type="image/x-icon" href="/wcsstore/ColesResponsiveStorefrontAssetStore/dist/30e70cfc76bf73d384beffa80ba6cbee/img/favicon.ico">
<link rel="stylesheet" href="/wcsstore/ColesResponsiveStorefrontAssetStore/dist/30e70cfc76bf73d384beffa80ba6cbee/css/google/fonts-Source-Sans-Pro.css" type="text/css" media="screen">

Tags anti-scrape, scrape protection

Challenge

How Imperva protects against scraping bots

Post author By admin
Post date November 6, 2020
No Comments on How Imperva protects against scraping bots

Imperva (that includes the former Distil anti-bot management) is a service providing many kinds of website protections. The present Imperva services include the following ones:

Cloud Web Application Firewall (WAF)
Bot Protection service (formerly Distil Networks)
IP Reputation Intelligence
Content Delivery Network (CDN)
Attack Analytics solution (eg. DDoS)

As to the protection of the bot scraping activities we mention the following.

Tags anti-scrape, scrape protection

Challenge

Scraping a Javascript-dependent website with puppeteer

Post author By admin
Post date June 25, 2020
No Comments on Scraping a Javascript-dependent website with puppeteer

Support us by purchasing the book (under $5) on this topic.

In today’s web 2.0 many business websites utilize JavaScript to protect their content from web scraping or any other undesired bot visits. In this article we share with you the theory and practical fulfillment of how to scrape js-dependent/js-protected websites.

Tags Javascript, Node.js, scrape protection

Development

Bypass Distil

The Distil scrape protection is a prominent one in the modern anti-scrape techniques. So, now we want to share with you some tips of how to bypass it. If you are interested, please make an inquiry to the following email: igor[dot]savinkin[at]gmail[dot]com

Tags anti-scrape, scrape protection