Save 5% every month: use code 5OFFSTORM at checkout
Home / Guides / LLM & RAG data collection

How to collect web data for LLMs and RAG through proxies

Set the proxy in the crawler, not the model: CrawlerRunConfig(proxy_config=...) in Crawl4AI, a proxies dict in requests, the proxy meta key in Scrapy. Run from an authorized IP, obey robots.txt, pace each domain, and deduplicate before you embed. Flat per-thread pricing keeps big recrawls predictable.

Updated October 2026AI agents & automation

Retrieval-augmented generation (RAG), fine-tuning sets and evaluation corpora all start the same way: fetch a lot of web pages, turn them into clean text or Markdown, split it into chunks and store it with its source. Tools like Crawl4AI and self-hosted Firecrawl were built for that last mile, returning LLM-ready Markdown instead of raw HTML, while Scrapy and requests still do the heavy lifting for large, static sites.

A proxy comes in once the crawl gets bigger than a few hundred pages. Sending every request from one IP gets you rate-limited, temporarily blocked or served degraded pages, and a corpus quietly full of “Access denied” and cookie-wall text is worse than a smaller clean one. Spreading requests across many IPs, at a polite pace, keeps the data you collect the data you meant to collect.

This guide covers the proxy settings for each tool, how many threads a browser-based crawler really uses, the rules that keep a crawl acceptable (robots.txt, site terms, rate limits), deduplication, and why billing per gigabyte is a poor match for rendered pages.

Which Storm plan fits

Dataset building is high-volume, mostly public pages, so rotating proxies fit: a new IP per connection on the Main gateway, a 700,000+ IP pool and unlimited bandwidth. Size the plan by how many pages you fetch at the same moment: an HTML-only crawler needs one thread per request, a headless-browser crawler about 10 per open page. Need a fixed IP for a source that allowlists you? Add a dedicated proxy.

Rotating proxies (from $14/mo): 700,000+ IPs behind fixed gateway IP:PORTs. New IP on every request, or every 3 or 15 minutes. USA, EU, USA+EU or Worldwide. Unlimited bandwidth on every plan.

Get 40 threads for $39/mo See all rotating proxies plans

Before you start

  1. Log in to the member area and copy your gateway IP:PORTs. They never change; the rotation happens on our side.
  2. Add the public IP of the computer or server that will run your tool under Authorized IPs, click Save, and allow up to 15 minutes before testing. Rotating and residential proxies use IP authentication, so there is no username or password.
  3. Dedicated proxies work with either IP authentication or a username and password. Use user:pass if your IP changes or the tool runs on several machines.
  4. Count your threads: the tool’s total open connections must stay within your plan (for example 40 threads on the 40-thread plan).

Set up a proxied crawl for an LLM pipeline

  1. Choose the fetcher per source

    Static pages (docs, blogs, forums, most news) only need an HTTP client: requests or Scrapy, which are fast and use one thread per request. Pages that build their content with JavaScript need a browser-based crawler such as Crawl4AI or Firecrawl’s Playwright service. Mixing both is normal: route only the sources that need rendering through the browser.

  2. Run it from a fixed, authorized IP

    Rotating gateways accept connections from IPs you save under Authorized IPs. Use your workstation or a VPS, and authorize its public IP (allow up to 15 minutes). Hosted notebooks and serverless jobs change IP between runs, so they can’t use IP authorization; run there only with a dedicated proxy and username:password.

  3. Add the proxy to the crawler

    In Crawl4AI, set proxy_config on CrawlerRunConfig; the old proxy argument on BrowserConfig is deprecated. One rotating gateway is enough: the rotation happens on the gateway, so you don’t need a list or a rotation strategy for it.

  4. Turn on robots.txt checks and identify your crawler

    Crawl4AI has check_robots_txt=True; Scrapy has ROBOTSTXT_OBEY = True. Set a user_agent that names your project and a contact URL, so a site owner can tell you apart from abuse and reach you.

  5. Cap concurrency to your threads

    For Crawl4AI, set the dispatcher’s max_session_permit to your threads divided by about 10 (40 threads = 4 pages at once). For Scrapy or requests, use your thread count or a little below it.

  6. Clean, deduplicate, then embed

    Drop error pages and boilerplate, normalise URLs, hash the cleaned text and skip anything you’ve already stored. Keep the source URL and fetch date with every chunk, so answers can cite it and you can refresh stale pages later.

Copy-paste config

Replace GATEWAY_IP:PORT with a rotating gateway from your member area. No username or password is needed for rotating gateways once your IP is authorized. For more on the HTTP clients, see the requests proxy guide and the Scrapy proxy guide.

Crawl4AI: one page to Markdown through the gateway
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, ProxyConfig

run_cfg = CrawlerRunConfig(
    proxy_config=ProxyConfig(server="http://GATEWAY_IP:PORT"),  # http:// for https sites too
    check_robots_txt=True,
    cache_mode=CacheMode.BYPASS,
    page_timeout=60000,
)
browser_cfg = BrowserConfig(headless=True, user_agent="MyRAGBot/1.0 (+https://example.com/bot)")

async def main():
    async with AsyncWebCrawler(config=browser_cfg) as crawler:
        result = await crawler.arun("https://httpbin.org/ip", config=run_cfg)
        print(result.markdown)

asyncio.run(main())
Crawl4AI: many URLs, capped to a 40-thread plan, with backoff
import asyncio
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, RateLimiter, ProxyConfig
from crawl4ai.async_dispatcher import SemaphoreDispatcher

THREADS = 40
PAGES_AT_ONCE = THREADS // 10          # a rendered page uses about 10 connections

run_cfg = CrawlerRunConfig(
    proxy_config=ProxyConfig(server="http://GATEWAY_IP:PORT"),
    check_robots_txt=True,
)
dispatcher = SemaphoreDispatcher(
    max_session_permit=PAGES_AT_ONCE,
    rate_limiter=RateLimiter(base_delay=(1.0, 3.0), max_delay=60.0,
                             max_retries=3, rate_limit_codes=[429, 503]),
)

async def main(urls):
    async with AsyncWebCrawler() as crawler:
        results = await crawler.arun_many(urls=urls, config=run_cfg, dispatcher=dispatcher)
        for r in results:
            print(r.url, r.success, len(r.markdown or ""))

asyncio.run(main(["https://example.com/a", "https://example.com/b"]))
Crawl4AI: dedicated proxy with username and password
from crawl4ai import CrawlerRunConfig, ProxyConfig

run_cfg = CrawlerRunConfig(
    proxy_config=ProxyConfig.from_string("PROXY_IP:PORT:USERNAME:PASSWORD")
)
Firecrawl self-hosted: proxy for the Playwright service (.env)
# apps/api/.env  (rebuild the containers after editing)
PROXY_SERVER=http://GATEWAY_IP:PORT
PROXY_USERNAME=
PROXY_PASSWORD=
# optional: skip images/video for faster pages
BLOCK_MEDIA=true
Deduplicate cleaned text before embedding
import hashlib, re

seen = set()

def fingerprint(text: str) -> str:
    norm = re.sub(r"\s+", " ", text).strip().lower()
    return hashlib.sha256(norm.encode()).hexdigest()

def keep(doc: dict) -> bool:
    """doc = {"url": ..., "markdown": ..., "fetched_at": ...}"""
    md = doc["markdown"] or ""
    if len(md) < 200 or "access denied" in md.lower():
        return False                     # error page or empty shell
    fp = fingerprint(md)
    if fp in seen:
        return False                     # same content under another URL
    seen.add(fp)
    return True

Threads: HTTP fetchers vs browser crawlers

The thread limit on a rotating plan counts open connections at one moment. How far that goes depends entirely on the fetcher:

  • requests, httpx, Scrapy: one connection per request in flight. A 40-thread plan runs about 40 parallel fetches, and a 150-thread plan about 150.
  • Crawl4AI, Firecrawl’s Playwright service, any headless browser: a rendered page opens many connections for scripts, styles, images and fonts, around 10 per page. The same 40-thread plan renders about 4 pages at once.

That’s why the cheapest way to scale a corpus is to render only what needs rendering. Many “JavaScript” sites still ship the text in the HTML or in a JSON endpoint the page calls; fetch that with an HTTP client and save the browser for the rest. Crawl4AI’s text_mode=True on BrowserConfig skips images and other heavy content, which makes pages lighter, but keep your page count based on 10 threads per page anyway.

Several crawlers sharing one plan share its threads. If a Scrapy job uses 30 of 40 threads, a Crawl4AI job next to it has room for one rendered page.

Collect responsibly: robots.txt, terms and rate limits

A proxy changes where requests come from. It doesn’t change what a site allows. For a dataset you want to keep using, follow a few rules:

  • robots.txt. Enable the robots check in your crawler. Many sites now list AI crawlers by name and disallow them; treat a disallow for your purpose as a no, even if your user agent string is different.
  • Terms of service. Some sites forbid automated collection or reuse of their content for model training. Read them for any source you rely on, and prefer sources with an explicit licence, an official dump or a public endpoint.
  • Per-domain pace. Concurrency should be spread across many domains, not aimed at one. A few requests at a time per site, with random delays and backoff on 429 and 503, is plenty for most sources.
  • No logins, no personal data. Stay on public pages. Don’t collect behind accounts, and filter personal data out of the text before it reaches an index.

Storm proxies may not be used for hacking or other illegal activity, and that includes collecting data a site has closed to you.

Deduplication and freshness

Duplicates cost twice: once in embedding and storage, and again at query time, when the retriever returns five near-identical chunks and pushes useful context out of the prompt. Deduplicate at three levels:

  • URL: strip tracking parameters and fragments, follow canonical tags, and keep a set of fetched URLs so a recrawl only fetches what’s new or changed.
  • Content: hash the cleaned Markdown (as in the code above) to catch mirrors, print versions and paginated copies.
  • Chunk: boilerplate such as cookie notices and footers repeats on every page; remove it before chunking, or hash chunks too.

Store fetched_at with every document. For refreshes, recrawl high-value sources on a schedule and compare hashes; unchanged pages don’t need new embeddings.

Per-GB pricing vs a flat plan for rendered pages

An LLM corpus is measured in pages, but most proxy plans are billed by data. A rendered page pulls in every script, font and image on it, so the bytes per page are many times the text you keep, and every retry, recrawl and refresh is billed again. Budgeting a big crawl on a per-GB plan means guessing page weight in advance and capping the crawler when the allowance runs out.

Storm rotating plans are priced by threads: 40 threads for $39, 80 for $59, 150 for $97 or 200 for $147 per month, with unlimited bandwidth. Recrawling the whole corpus every week costs the same as crawling it once. The number that matters is how many pages you fetch in parallel, which you control in the crawler settings.

Where to run the crawler

Because rotating gateways authorize by IP, the crawler should live somewhere with a stable public IP: your own machine, an office server, or a small VPS. Hosted notebooks, serverless functions and many CI runners get a different IP each run, so the gateway won’t recognise them.

A practical layout is a VPS that runs the crawler on a schedule and writes Markdown plus metadata to storage, with embedding done wherever your vector database lives. If part of the pipeline must run in the cloud with a changing IP, give that part a dedicated proxy with username and password. If you also have AI agents browsing pages interactively, see the AI browser agent guide.

Common errors and fixes

Crawl4AI ignores the proxy, your IP showsThe proxy was set with the deprecated proxy argument on BrowserConfig, or the config isn’t passed to arun. Put proxy_config on CrawlerRunConfig and pass that config with every call.
Proxy or tunnel connection error / connection reset on every pageThe crawler’s machine isn’t on your Authorized IPs list, or it was added less than 15 minutes ago; the gateway accepts the connection and resets it. On a VPS, authorize the VPS IP. A plain “connection refused” means a wrong port or a firewall instead.
SSL or tunnel errors on every HTTPS pageThe proxy URL starts with https://. Use http://GATEWAY_IP:PORT; pages still load over HTTPS through the tunnel.
Lots of timeouts once the crawl scales upToo many rendered pages for your threads. Divide threads by about 10 to get pages at once, lower max_session_permit, and raise page_timeout for slow sites.
Corpus full of “Access denied”, CAPTCHA or cookie-wall textYour pace per domain is too high, or the site doesn’t allow crawling. Slow down, back off on 429/503, check robots.txt, and filter error pages out before indexing.
Firecrawl self-host still uses your IPThe .env was edited after the containers started. Rebuild and restart them so the Playwright service reads PROXY_SERVER.

FAQ

What is the best proxy type for LLM training data collection?

For public pages at scale, rotating datacenter proxies: a large pool, a new IP per connection, and pricing that doesn’t grow with data. Use static dedicated proxies only for sources that need a fixed, allowlisted IP.

How do I set a proxy in Crawl4AI?

Pass proxy_config=ProxyConfig(server="http://GATEWAY_IP:PORT") (or a plain string) to CrawlerRunConfig. For a list of proxies, Crawl4AI has RoundRobinProxyStrategy, but a rotating gateway already rotates for you.

How many pages at once can I crawl on a Storm 40-thread plan?

About 40 with an HTTP client like Scrapy or requests, or about 4 rendered pages with a headless-browser crawler such as Crawl4AI, because each rendered page uses around 10 threads.

Do Storm rotating proxies work from Google Colab or a serverless function?

Usually not. Rotating plans use IP authorization only, and those environments change IP between runs. Run the crawler on a machine with a fixed IP, or use a dedicated proxy with username and password there.

Is it legal to scrape websites for RAG or AI training?

It depends on the site’s terms, the content’s copyright and your jurisdiction, and this guide isn’t legal advice. Respect robots.txt and terms, stay on public pages, keep personal data out, and prefer licensed sources or official dumps where they exist.

Does Storm charge for bandwidth on large crawls?

No. Every plan has unlimited bandwidth and is priced per thread, port or proxy, so full-page renders and repeated recrawls don’t change the bill.

Still have questions? Contact us here. A real person answers.

Related guides

Tool facts checked against the official documentation (October 2026): Crawl4AI: proxy and security · Crawl4AI: parameters reference · Crawl4AI: multi-URL crawling and dispatchers · Firecrawl: apps/api/.env.example · Scrapy settings: ROBOTSTXT_OBEY. Storm Proxies facts: our plans page and refund policy.

Unlimited bandwidth. One flat monthly price.

Access is live the moment you pay, and the smallest package of each proxy type has a 24-hour money-back guarantee on your first order.