Save 5% every month: use code 5OFFSTORM at checkout
Home / Guides / Scrapy

How to set up a proxy in Scrapy

Set request.meta['proxy'] = 'http://GATEWAY_IP:PORT' on each request, or do it once in a small downloader middleware. Scrapy’s built-in HttpProxyMiddleware picks it up for both http and https URLs. Then cap CONCURRENT_REQUESTS at your plan’s thread count so the crawl never opens more connections than you pay for.

Updated October 2026Code & scraping frameworks

Scrapy is the Python framework people pick when a script with requests stops being enough: it schedules thousands of URLs, runs them concurrently on an event loop, follows links, retries failures and exports items. Proxy support is already inside it. A downloader middleware called HttpProxyMiddleware is enabled by default and routes any request that carries a proxy key in its meta.

The setup is short. The trouble usually starts later: the spider runs 16 requests at once by default (or far more after someone raises the setting), a 10-thread test plan can’t keep up, retries pile onto a site that’s already rate-limiting, and a spider that should rotate keeps one IP because Scrapy holds connections open. Each of those has a clean fix in settings.

Below you’ll find four ways to attach a proxy, a reusable middleware for one or many gateways, the settings that match Scrapy’s concurrency to your plan, how retries and AutoThrottle behave behind a rotating gateway, and how to keep the same setup when you add scrapy-playwright for JavaScript pages.

Which Storm plan fits

Scrapy is a fast, connection-hungry crawler, so rotating proxies are the natural match: one Main gateway spreads requests over a large pool, and the thread count maps straight onto CONCURRENT_REQUESTS. If your spider logs in and must keep one identity, send those requests through the 3- or 15-minute gateway, a residential port or a private dedicated proxy instead.

Rotating proxies (from $14/mo): 700,000+ IPs behind fixed gateway IP:PORTs. New IP on every request, or every 3 or 15 minutes. USA, EU, USA+EU or Worldwide. Unlimited bandwidth on every plan.

Get 40 threads for $39/mo See all rotating proxies plans

Before you start

  1. Log in to the member area and copy your gateway IP:PORTs. They never change; the rotation happens on our side.
  2. Add the public IP of the computer or server that will run your tool under Authorized IPs, click Save, and allow up to 15 minutes before testing. Rotating and residential proxies use IP authentication, so there is no username or password.
  3. Dedicated proxies work with either IP authentication or a username and password. Use user:pass if your IP changes or the tool runs on several machines.
  4. Count your threads: the tool’s total open connections must stay within your plan (for example 40 threads on the 40-thread plan).

Add a proxy to a Scrapy project

  1. Create or open the project

    Run pip install scrapy, then scrapy startproject shop and scrapy genspider products example.com. Everything below goes in settings.py, middlewares.py or the spider file.

  2. Pick a gateway for the job

    From the member area, copy the gateway you need. Crawling catalog or search pages: the Main rotating gateway. A session that logs in or fills a cart: the 15-minute gateway or a residential port, so the IP holds while the session lasts.

  3. Attach the proxy to requests

    The fastest test is in the spider: yield scrapy.Request(url, meta={'proxy': 'http://GATEWAY_IP:PORT'}). For a whole project, set it in a downloader middleware (code below) so no spider can forget it. The value always starts with http://, even when the target URL is https.

  4. Match concurrency to your plan

    In settings.py set CONCURRENT_REQUESTS to your thread count or a little under it. On a 40-thread plan, 36–40 is right. Leave headroom if a browser or a second spider uses the same plan.

  5. Confirm the exit IP

    Run the spider in example 1, or open scrapy shell and fetch() a Request for https://httpbin.org/ip with the proxy in its meta. You should see a pool IP. Your own IP means the proxy key never reached the request; a connection that’s reset or lost right away means your machine’s IP isn’t authorized yet.

  6. Turn on retries and throttling

    Keep RetryMiddleware on, add 429 handling and enable AutoThrottle for sites that slow down under load. Then start the crawl and watch the log’s stats at the end for retry/count and response codes.

Scrapy proxy code you can paste

Swap GATEWAY_IP:PORT for a gateway from your member area. Rotating and residential gateways authenticate by IP, so the proxy value has no username or password. Only private dedicated proxies use the user:pass@ form.

1. Proxy on a single request (spider)
import scrapy

PROXY = "http://GATEWAY_IP:PORT"

class IpSpider(scrapy.Spider):
    name = "ip"

    async def start(self):                 # Scrapy 2.13+; use start_requests() on older versions
        for i in range(5):
            yield scrapy.Request("https://httpbin.org/ip", meta={"proxy": PROXY},
                                 dont_filter=True, cb_kwargs={"n": i})

    def parse(self, response, n):
        self.logger.info("request %s left from %s", n, response.json()["origin"])
2. Project-wide middleware (middlewares.py)
import random

class StormProxyMiddleware:
    """Puts a gateway on every request that doesn't already have one."""

    def __init__(self, gateways):
        self.gateways = gateways

    @classmethod
    def from_crawler(cls, crawler):
        gws = crawler.settings.getlist("STORM_GATEWAYS")
        if not gws:
            raise ValueError("Set STORM_GATEWAYS in settings.py")
        return cls(gws)

    def process_request(self, request, spider):
        if "proxy" not in request.meta:            # a spider can still override it
            request.meta["proxy"] = random.choice(self.gateways)
3. settings.py for a 40-thread rotating plan
STORM_GATEWAYS = [
    "http://GATEWAY_IP:PORT",       # Main gateway (new IP per connection)
    # add more gateway ports from your plan here
]

DOWNLOADER_MIDDLEWARES = {
    "shop.middlewares.StormProxyMiddleware": 740,   # runs before HttpProxyMiddleware (750)
}

CONCURRENT_REQUESTS = 38            # total open requests, keep <= plan threads
CONCURRENT_REQUESTS_PER_DOMAIN = 8  # per-site cap; raise only if the site copes
DOWNLOAD_DELAY = 0.5
DOWNLOAD_TIMEOUT = 40               # default is 180 seconds

RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429]

AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1.0
AUTOTHROTTLE_MAX_DELAY = 30.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 4.0

USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0 Safari/537.36"
4. Environment variables (no code changes)
# macOS / Linux
export http_proxy="http://GATEWAY_IP:PORT"
export https_proxy="http://GATEWAY_IP:PORT"
scrapy crawl products

# A meta['proxy'] value still wins over these variables.
5. Private dedicated proxy with username and password
from urllib.parse import quote

PWD = quote("PASSWORD", safe="")    # encode @ : / and similar characters
yield scrapy.Request(url, meta={"proxy": f"http://USERNAME:{PWD}@PROXY_IP:PORT"})
# HttpProxyMiddleware turns the credentials into a Proxy-Authorization header

CONCURRENT_REQUESTS is your thread count

Storm plans limit how many connections are open at the same moment. Scrapy’s matching knob is CONCURRENT_REQUESTS: the most requests the downloader handles at once across every domain. Its built-in default is 16. That already overloads the 10-thread test package and leaves most of a 150-thread plan unused.

  • 10 threads (test package): CONCURRENT_REQUESTS = 8. Enough to prove the spider works, not to run it in production.
  • 40 threads: 36–40. 80 threads: 72–80. 150 threads: 135–150. 200 threads: 180–200.
  • Two spiders on one plan: split the number. Two spiders at 40 each on a 40-thread plan means 80 connections against a 40 limit, and roughly half of them fail.
  • Search engine result pages: no more than 25% of the plan, through the Main gateway. A 40-thread plan gets CONCURRENT_REQUESTS = 10 for that spider.

Then set CONCURRENT_REQUESTS_PER_DOMAIN for politeness. Projects created with a recent scrapy startproject ship with a per-domain value of 1 and a 1-second delay, which is why a fresh single-site spider can feel slow even on a big plan. Raise those two deliberately, per site, rather than only cranking the global number.

Why a Scrapy spider can keep one IP on a rotating gateway

The Main gateway assigns a new exit IP when a connection or HTTPS tunnel is opened. Scrapy’s HTTP/1.1 handler keeps a pool of persistent connections and reuses them, so a burst of requests to one site may travel through the same tunnel and share an IP. Usually that’s harmless; the pool spreads load over many connections anyway.

If a test with five quick requests to httpbin shows the same IP twice, that’s connection reuse, not a stuck gateway. Spread the same test over several domains or more time and the IPs vary. When you need the opposite, the same IP across a login and the pages behind it, don’t fight the Main gateway. Route that spider through the 3- or 15-minute gateway, or a residential port, where the IP holds by design. Note that the 3-minute gateway is meant for sign-ups, social sites and browsing, not for search engine scraping.

Retries and AutoThrottle behind a rotating gateway

RetryMiddleware is on by default and retries twice for codes 500, 502, 503, 504, 522, 524, 408 and 429, plus timeouts and dropped connections. Behind a rotating gateway a retry is useful: whenever it goes out on a new connection, it leaves from a different IP. The default RETRY_PRIORITY_ADJUST of -1 also schedules a retry behind fresh requests, which gives a rate-limited site a short break before it sees the URL again. Three retries is a sensible ceiling. Beyond that you’re usually hammering a page that’s gone or a site that wants you to slow down.

AutoThrottle adjusts the delay from measured latency. It starts at AUTOTHROTTLE_START_DELAY (5 seconds by default), never goes below DOWNLOAD_DELAY, and aims for AUTOTHROTTLE_TARGET_CONCURRENCY parallel requests per site (1.0 by default). Error responses can raise the delay but never lower it, which is exactly what you want when a site starts returning 503s. Two things to remember: AutoThrottle works per site, not across your plan, so CONCURRENT_REQUESTS still has to respect your thread limit; and the per-domain cap still applies, so a target concurrency above it is never reached.

One gateway, several gateways, or per-spider gateways

A gateway IP:PORT never changes, and the rotation happens on our side. So in most projects you don’t need a proxy list, a rotation package or a “ban detection” plugin at all. One Main gateway already sends traffic out through the whole pool.

Packages like scrapy-rotating-proxies exist for lists of static proxies, where the code itself must take a dead proxy out of rotation. You only need that pattern with dedicated proxies: give the middleware above your list of dedicated IP:PORTs and it spreads requests over them. For residential, list one port per entry; each port is a separate IP, so 10 ports in the list means requests leave from 10 IPs at the same moment.

Mixing gateways inside one project is common: catalog spiders on the Main gateway, an account spider on the 15-minute one. Set custom_settings or a class attribute per spider, and have the middleware read it instead of the global list.

JavaScript pages: scrapy-playwright and threads

When a page builds its content in the browser, add scrapy-playwright. Its download handler ignores meta['proxy'], which is a documented limitation. Set the proxy in PLAYWRIGHT_LAUNCH_OPTIONS for the whole browser, or define named contexts in PLAYWRIGHT_CONTEXTS, each with its own proxy, and pick one per request with meta={'playwright': True, 'playwright_context': 'eu'}.

Budget threads differently here. A rendered page loads scripts, images and API calls in parallel, around 10 connections per open page. Cap pages with PLAYWRIGHT_MAX_PAGES_PER_CONTEXT and PLAYWRIGHT_MAX_CONTEXTS so that pages × 10 stays inside your plan. On a 40-thread plan that’s about 4 pages, or fewer if plain Scrapy requests share the plan. The Playwright guide covers browser-level setup in more depth.

Common errors and fixes

Connection to the other side was lost / ConnectionResetErrorThe machine running Scrapy isn’t on your Authorized IPs, or the save is less than 15 minutes old. The gateway accepts the connection and resets it. Check your public IP, compare it with the member area and wait out the 15 minutes.
TCP connection timed out / Connection was refused by other sideNothing answered on that address: the gateway IP or port is mistyped, or a firewall blocks the port. Copy the gateway again from the member area.
Could not open CONNECT tunnel with proxyThe HTTPS tunnel wasn’t opened: usually authorization (the gateway closed the connection), or more open connections than your plan allows. Lower CONCURRENT_REQUESTS and confirm the IP is authorized.
407 Proxy Authentication RequiredOn dedicated proxies with user:pass, the credentials are wrong or a special character isn’t URL-encoded. On rotating and residential an unauthorized IP shows as a reset, not a 407, so a 407 there comes from another proxy in the chain. See the 407 guide.
Spider returns your real IPNo request carried the proxy: the middleware isn’t in DOWNLOADER_MIDDLEWARES, its path is wrong, or it’s ordered after 750. Check the startup log, which lists every enabled middleware.
Many 429 responses or retry/max_reached in statsToo many parallel requests to one site. Lower CONCURRENT_REQUESTS_PER_DOMAIN, enable AutoThrottle and give the delay room. More retries won’t fix it. See fixing 429.
scrapy-playwright pages show your real IPIts handler ignores meta['proxy']. Put the proxy in PLAYWRIGHT_LAUNCH_OPTIONS or in a context in PLAYWRIGHT_CONTEXTS.

FAQ

How do I rotate proxies in Scrapy with Storm Proxies?

You don’t need a rotation package. Point every request at the Main rotating gateway and the exit IP changes on our side for each new connection. Use a list in the middleware only when you combine several gateway ports or a set of dedicated proxies.

What should CONCURRENT_REQUESTS be on a Storm rotating plan?

Your thread count or slightly below: about 38 on the 40-thread plan, 75 on the 80-thread plan. For search engines use no more than a quarter of the plan. All spiders and tools on the same plan share that limit.

Does Scrapy support HTTPS through an HTTP proxy?

Yes. For https URLs Scrapy opens a CONNECT tunnel through the proxy and runs TLS inside it. Write the proxy as http://GATEWAY_IP:PORT; the target URL stays https.

Is meta['proxy'] or the http_proxy environment variable better?

Use meta['proxy'] (set by a middleware) for real projects. It’s explicit, can differ per spider, and wins over the environment variables. The variables are handy for a quick test or for running someone else’s spider unchanged.

Can I set a proxy per spider?

Yes. Give each spider a proxy_gateways attribute or a value in custom_settings, and let your middleware read it in process_request. That keeps a login spider on a 15-minute gateway while catalog spiders use the Main one.

Can I use Storm proxies on Scrapy Cloud?

Only with a setup that doesn’t depend on a fixed source IP. Rotating and residential gateways need an authorized IP, and hosted jobs usually change IP between runs. Private dedicated proxies with username and password work there; or run the spider on your own VPS.

Still have questions? Contact us here. A real person answers.

Related guides

Tool facts checked against the official documentation (October 2026): Scrapy: Downloader middleware (HttpProxyMiddleware, RetryMiddleware) · Scrapy: Settings · Scrapy: AutoThrottle · scrapy-playwright README. Storm Proxies facts: our plans page and refund policy.

Unlimited bandwidth. One flat monthly price.

Access is live the moment you pay, and the smallest package of each proxy type has a 24-hour money-back guarantee on your first order.