![]()
How To Customise Scrapy: Extensions, Middlewares & Pipelines Explained
What makes Scrapy great, isn't just the fact that it comes with so much functionality out of the box, it is because Scrapy's core functionality is so easily customizable when you understand how Scrapy works and how you can create your own Scrapy Extensions, Downloader Middlewares, and Spider Middlewares.
With over 150+ of open-source Scrapy extensions available and the ability to easily create your own customizations, you can greatly improve Scrapy's performance to your particular use case with little to no effort.
In this guide, we're going to explain how to customize Scrapy using Extensions, Middlewares and Pipelines.
- Why Customize Scrapy?
- Extensions, Middlewares, and Pipelines: An Overview and Comparison
- Understanding Scrapy Extensions
- Scrapy Middlewares: Customizing the Spider Workflow
- Scrapy Pipelines: Managing Item Processing
- Integrating Custom Components in Scrapy
- Advanced Tips for Customizing Scrapy
- Real-World Examples of Scrapy Customization
- Conclusion
- More Python Web Scraping Guides
Why Customize Scrapy?
Scrapy is a full fledged and fully customizable web scraping framework. When it comes to performance, almost nothing beats Scrapy.
It's asynchronous by default which allows you to crawl multiple pages concurrently right out of the box! Scrapy is arguably the most efficient Python scraping framework out there.
However, when you start a new Scrapy project, it's pretty barebones. Every scrape is different and the developers behind Scrapy understand this. When you customize your Scrapy project, you can tweak it to better fit your target site and your desired results.
With custom extensions, you can alter the global behavior of your project. If you want to save parts of your log to a file, we can do that in well under 50 lines of code.
With custom middleware, you can alter the behavior of your scraper during operation to add things like retry logic.
With custom pipelines, Scrapy allows you to do whatever you want with your data, whether you want to save it to a CSV, or update a database directly!
Before continuing, you should create a new Scrapy project. You can use it to follow along and customize your own Scrapy project.
scrapy startproject my_custom_crawler
Extensions, Middlewares, and Pipelines: An Overview and Comparison
In Scrapy, extensions, middlewares, and pipelines play distinct yet interconnected roles in customizing and optimizing your web scraping workflow.
- Extensions extend Scrapy’s core functionality, enabling tasks like monitoring performance or sending alerts.
- Middlewares act as intermediaries between Scrapy’s engine and requests or responses, making them ideal for handling user-agent rotation, proxy management, or request modifications.
- Finally, pipelines focus on processing and storing scraped data, handling tasks like cleaning, validating, and exporting items.
To help you better understand their roles and how they differ, we’ve included a comparison table that outlines their key functions, typical use cases, and implementation details.
This table will serve as a quick reference to guide you in determining when and how to use each component effectively in your projects.
| Aspect | Extensions | Middleware | Pipelines |
|---|---|---|---|
| Primary Purpose | Modify global project behavior | Customize request and response processing | Process and store scraped data |
| Scope | Project-wide (affects all spiders and processes) | Operates on requests (Downloader Middleware) or responses (Spider Middleware) during the crawl | Focused on individual scraped items |
| Common Use Cases | - Logging - Monitoring - Error reporting - Alerts | - Retry logic - Proxy integration - Adding custom headers - Response modification | - Data cleaning - Validation - Storing data (e.g., saving to database or file) |
| Integration Level | High-level hooks triggered by Scrapy signals | Middle-layer between Scrapy engine and Downloader/Spider | Post-processing step for handling extracted items |
| Advantages | - Centralized modifications - Easy to monitor spider activity | - Flexibility for handling requests/responses - Granular control over scraping logic | - Simplifies data processing and storage logic |
| Disadvantages | - Errors can impact the entire project - Requires careful testing | - Can introduce complexity in managing requests/responses - Needs configuration for specific tasks | - Can be resource-intensive with large datasets - Requires proper validation to avoid errors |
Understanding Scrapy Extensions
Scrapy extensions provide a powerful way to augment the functionality of your scraping projects, allowing you to customize and enhance Scrapy’s behavior without modifying its core framework.
Extensions are particularly useful for tasks such as monitoring the performance of your spiders, logging important events, sending alerts, or integrating with external tools and services.
What are Scrapy Extensions?
Extensions are used to modify the global behavior of your project. Whether you run multiple crawls using multiple spiders, or just one single spider, your project will always follow the rules set by your extensions.
If we create an extension to log the final status of each response, every time we're finished with a response, the status will come up in our log.
- Scope: Extensions modify global behavior of the project. If you implement a logging extension, it will trigger whenever a spider is open or a crawl has started.
- Usecases: Extensions are great for monitoring your entire Scrapy system. Think of the logging example again. If you receive a bad HTTP response, this will be recorded in the log for you to inspect later. If you want to receive an email when your scraper encounters an error, you'd use an extension to implement this.
- Pros: This greatly enhances the flexibility of the full project. With a relatively small amount of code, you can make large changes to the project in a single place.
- Cons: Because extensions use a global scope, it can lead to global bugs if you implement something wrong. Because of their global nature, extensions shouldn't interact with our request/response protocol or our extraction logic.
Built-in vs. Custom Extensions
There are many extensions that come built-in with Scrapy. These extensions are useful tools for debugging, pausing, throttling, memory management and much more.
You can get a better understanding of these extensions below.
- TelnetConsole: This allows you to debug, inspect and modify your scraper during runtime.
- LogStats: This extension logs basic stats from your scraper such as requests and pages you've scraped.
- CoreStats: Collect core stats such as item counts, status code counts, and saves them into the stats collector.
- AutoThrottle: Throttles your scraper based on server responses. This is incredibly useful in managing latency.
- MemoryUsage: Monitor the memory usage of your scraper. You can set a shutoff when your memory exceeds a specific usage level.
Limitations of Built-in Extensions
Built-in extensions are sort of like a "one size fits all" piece of clothing. They're meant to cover basic usage for the average person.
If you want to perform specific actions, these built-in extensions limit what you're able to do because they're meant to be generic.
There's also a bit of performance overhead with these extensions because they're made to do so much.
If you're using an extension for one piece of functionality, you don't need the other 5 functions that come with it.
Popular Open Source Scrapy Extensions
The Scrapy ecosystem boasts a variety of open-source extensions that enhance its functionality and simplify common scraping tasks.
Here are some popular open-source Scrapy extensions:
- ScrapeOps Scrapy Extension: This one comes right from us.
- Spidermon: Monitor your spiders to check your output data and create custom alerts. Get custom notifications via Slack, Telegram, Discord and email.
- Spider Feeder: Place a file inside your Scrapy project and crawl every url from the file. Use almost any file type:
.txt,.csv,jsonquickly without the need for boilerplate code. - Scrapy Statsd: Send your stats to a hosted server. This can help you to manage many scrapers all from one place.
- Scrapy JSONRPC: Control your scraper using JSON RPC commands.
Steps to Create a Custom Extension
When creating an extension, we need to follow several steps.
- We need to identify the hook for the extension.
- Then, we write the extension class.
- Finally, we adjust our Scrapy settings to use the new extension.
1. Identify the Hook
There are a number of hooks we can use to trigger our extension. The table below outlines a few of the most common ones.
| Hook | Triggered When | Use Case |
|---|---|---|
spider_opened | A spider starts crawling. | Initialize resources |
spider_closed | A spider finishes crawling. | Cleanup resources/Create a report. |
response_received | A response is received. | Process/Log response data globally. |
engine_started | The Scrapy engine starts running. | Setup global monitoring systems. |
item_scraped | An item is successfully scraped. | Track or process scraped data. |
request_dropped | Request has been dropped/filtered. | Log dropped requests for debugging. |
2. Write the Extension Class
Next, you need to write your extension class. Here's an extension for a custom logger that we'll use.
class CustomLoggerExtension:
def __init__(self, stats, log_file):
self.stats = stats
self.log_file = log_file
self.start_time = None
self.end_time = None
self.logger = logging.getLogger("custom_logger")
self.logger.setLevel(logging.INFO)
handler = logging.FileHandler(self.log_file)
handler.setLevel(logging.INFO)
formatter = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s - %(message)s')
handler.setFormatter(formatter)
self.logger.addHandler(handler)
@classmethod
def from_crawler(cls, crawler):
log_file = crawler.settings.get("CUSTOM_LOG_FILE", "custom_log.txt")
ext = cls(crawler.stats, log_file)
crawler.signals.connect(ext.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(ext.spider_closed, signal=signals.spider_closed)
crawler.signals.connect(ext.response_received, signal=signals.response_received)
return ext
def spider_opened(self, spider):
self.logger.info(f"Spider {spider.name} opened.")
self.start_time = datetime.now()
def spider_closed(self, spider):
self.end_time = datetime.now()
runtime = self.end_time - self.start_time
self.logger.info(f"Spider {spider.name} closed. Time elapsed: {runtime} seconds")
def response_received(self, response, request, spider):
self.logger.info(f"Response received, status: {response.status} for {response.url}")
3. Activate the Extension
Once you've created your class, you need to adjust your settings to account for this new class.
Scroll down to the EXTENSIONS section of settings.py.
Uncomment the section to turn on custom extensions and add the path to the extension you just created.
# Enable or disable extensions
# See https://docs.scrapy.org/en/latest/topics/extensions.html
EXTENSIONS = {
"scrapy.extensions.telnet.TelnetConsole": None,
"my_custom_crawler.extensions.custom_logger.CustomLoggerExtension": 500
}
Example: Creating an Extension for Runtime Stats Collection
The following extension monitors some basic stats for us and saves them to a .log file.
- Whenever a crawler is opened, we get the current time with the
datetimemodule. - Once the spider is closed, we get the
datetimeagain. - We use these times to calculate the total runtime of the spider.
- We also log the status code of any responses that the spider receives.
Here is the full code to our custom logger.
import logging
from datetime import datetime
from scrapy import signals
class CustomLoggerExtension:
def __init__(self, stats, log_file):
self.stats = stats
self.log_file = log_file
self.start_time = None
self.end_time = None
self.logger = logging.getLogger("custom_logger")
self.logger.setLevel(logging.INFO)
handler = logging.FileHandler(self.log_file)
handler.setLevel(logging.INFO)
formatter = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s - %(message)s')
handler.setFormatter(formatter)
self.logger.addHandler(handler)
@classmethod
def from_crawler(cls, crawler):
log_file = crawler.settings.get("CUSTOM_LOG_FILE", "custom_log.txt")
ext = cls(crawler.stats, log_file)
crawler.signals.connect(ext.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(ext.spider_closed, signal=signals.spider_closed)
crawler.signals.connect(ext.response_received, signal=signals.response_received)
return ext
def spider_opened(self, spider):
self.logger.info(f"Spider {spider.name} opened.")
self.start_time = datetime.now()
def spider_closed(self, spider):
self.end_time = datetime.now()
runtime = self.end_time - self.start_time
self.logger.info(f"Spider {spider.name} closed. Time elapsed: {runtime} seconds")
def response_received(self, response, request, spider):
self.logger.info(f"Response received, status: {response.status} for {response.url}")
from_crawler(): This is essentially the runtime for the extension. This triggers every time a spider is opened.spider_opened(): This saves the current time when the spider is opened.spider_closed(): Checks the time again and calculates the total runtime of the spider.
Scrapy Middlewares: Customizing the Spider Workflow
Middlewares in Scrapy serve as powerful intermediaries that allow you to customize and control the flow of requests and responses between the Scrapy engine and your spiders.