Back to blog

Guide

How to Prevent Request Throttling at Scale

Learn how to prevent request throttling with smarter pacing, retries, caching, and proxy rotation for reliable, scalable web data collection at scale.

Throttling is the target's response to a request pattern it considers excessive. The instinct when it happens is to add more IPs, but that treats the symptom. Most throttling is caused by a request pattern that is easy to fix, and the fixes compound: pace correctly, retry correctly, stop sending requests you do not need, and distribute what remains across enough IPs that no single one looks heavy. In that order.

Understand What Triggers It

Targets throttle on observed patterns, not on intent. The common triggers:

Per-IP request volume in a window. The most basic form. Exceed N requests per minute from one IP and you get rate-limited.

Burst shape. Fifty requests in two seconds followed by silence reads as automated far more clearly than fifty requests spread across a minute, even though the volume is identical.

Request regularity. Perfectly even intervals — exactly 500ms apart — are a machine signature. Real traffic is irregular.

Subnet and ASN aggregation. Many targets aggregate at the /24 or ASN level, so rotating within a narrow address range does not distribute the load as much as the IP count suggests.

Concurrent connection count per IP. Opening twenty simultaneous connections from one IP is itself a signal, independent of total volume.

Navigation coherence. Requesting deep pages with no referrer, no asset loads, and no session history looks unlike a browsing user.

Each of these is addressable. Only the first is addressed by adding IPs.

Pace Deliberately

Pacing is the highest-leverage fix and the one most often skipped.

Set an explicit rate per target. Not a global rate — per target, derived from measured tolerance. Leaving concurrency uncapped and letting the job run at maximum throughput is how most throttling starts.

Jitter every interval. Replace fixed delays with randomized ones. A uniform or normal distribution around your target interval removes the regularity signature at essentially no cost:

import random, time
 
def paced_sleep(base_seconds, jitter=0.4):
    # e.g. base 2.0s with 0.4 jitter -> 1.2s to 2.8s
    delay = base_seconds * (1 + random.uniform(-jitter, jitter))
    time.sleep(max(0.1, delay))

Smooth bursts. If your scheduler produces work in batches, spread each batch across the interval rather than firing it all at once. A token-bucket or leaky-bucket limiter per target does this cleanly and is worth the small amount of code.

Cap concurrent connections per IP. Keep it low — one or two for most targets. High per-IP concurrency is a strong signal and buys you little, since the right way to increase throughput is more IPs, not more connections per IP.

Retry Correctly

Bad retry logic turns a transient problem into a sustained one. The rules:

Never retry a 429 or 403 on the same IP. The IP has been flagged. Retrying on it adds to the pattern that caused the flag and can escalate a rate limit into a longer block. Rotate to a new session first.

Honor Retry-After. When the target tells you when to come back, comply. Ignoring it is the fastest way to convert throttling into a ban.

Use exponential backoff with jitter. Fixed-interval retries from many workers synchronize into bursts. Add randomness:

def backoff_delay(attempt, base=1.0, cap=60.0):
    raw = min(cap, base * (2 ** attempt))
    return raw * random.uniform(0.5, 1.0)   # decorrelate concurrent retriers

Retry 5xx on the same IP. Target server errors are not the IP's fault. Rotating wastes a session start for a problem rotation does not fix.

Cap total attempts and record the failure. Infinite retry loops against a blocking target generate load, consume bandwidth, and accomplish nothing. Three attempts is usually right; log the permanent failure and move on.

Add a circuit breaker per target. When a target's failure rate crosses a threshold, stop sending to it for a cooldown period rather than continuing to hammer it. This is the single most effective protection against an incident where a target tightens its limits and your job keeps pushing into the wall.

Send Fewer Requests

The requests you never send cannot be throttled. This is consistently the most overlooked lever.

Cache aggressively. Respect ETag and Last-Modified, and send conditional requests. A 304 response costs almost nothing in bandwidth and does not count against most rate limits the way a full fetch does.

Deduplicate before dispatch. Jobs that enqueue the same URL from multiple paths waste requests. Dedupe the queue.

Prefer structured endpoints over page scraping. If a target exposes a public JSON endpoint serving the same data, one API call can replace a full page load plus assets — often a 10× reduction in requests and far more in bandwidth.

Skip unnecessary assets. When using browser automation, block images, fonts, media, and analytics unless you need them. A product page might be 40 requests with assets and 3 without. This cuts both your throttling exposure and your bandwidth bill substantially:

await page.route("**/*", lambda route: (
    route.abort()
    if route.request.resource_type in ("image", "font", "media", "stylesheet")
    else route.continue_()
))

Use plain HTTP where JavaScript is not required. Full browser rendering multiplies request count. Reserve it for targets that genuinely need it.

Collect only what changes. Prioritize by volatility. Refreshing stable data on the same cadence as volatile data wastes most of the requests.

Distribute Across IPs Properly

Once pacing and volume are right, distribution is what lets you scale.

Many IPs, each lightly used. This is the correct shape: wide parallelism where every individual IP stays well under the target's tolerance. It is the opposite of pushing a small pool hard.

Rotate on a request budget, not only on failure. Retire a session proactively after N requests against a target — below measured tolerance — rather than waiting for a block. Rotating after failure means you have already taken the hit.

Spread across ASNs. Since many targets aggregate by network, diversity of ASN matters as much as count of IPs. Monitor your per-ASN distribution and avoid concentrating.

Give IPs rest. An IP that sits idle for a few hours partially recovers its budget against a target. Drawing from a large pool with long rest intervals keeps per-IP history low.

Isolate sessions per target. An IP with heavy history against one target starts fresh against another. Keep session namespaces separate so you are not carrying reputation across targets unnecessarily.

Look Like a Browser

Throttling often triggers on coherence rather than volume alone.

Set a realistic, current User-Agent. Send the full header set a real browser sends, in a plausible order. Include Referer when the navigation implies one. Match Accept-Language and timezone to the proxy's geography. For browser automation, use stealth configuration and realistic viewport and device settings. Maintain cookies within a session rather than discarding them on every request — a session that never carries state looks unlike a user.

None of this is about evasion; it is about not emitting signals that misrepresent legitimate collection as something else.

The Throughput Trap

Here is where providers matter. Everything above points toward the same architecture: many concurrent sessions, each paced gently. That is how you reach volume without throttling.

A provider that caps concurrent sessions forces the opposite. If you need 100,000 requests in a window and your thread allowance is 100, each of those 100 IPs has to carry 1,000 requests — far above most targets' tolerance. The provider's cap pushes you into exactly the per-IP pattern that triggers throttling, and no amount of good pacing logic fixes it. You end up choosing between getting throttled and not finishing the job.

FlameProxies applies no concurrency cap and no request-rate limits, so you can run the wide-and-gentle pattern that actually prevents throttling: thousands of parallel sessions each sending a handful of well-paced requests. With 80M+ residential IPs across 180+ countries and city-level targeting, the distribution and rest intervals that keep per-IP history low are available rather than rationed.

Checklist

  • Explicit rate limit per target, derived from measured tolerance
  • Jitter on all intervals; no fixed delays
  • Per-IP concurrent connections capped at 1–2
  • Never retry 429/403 on the same IP
  • Retry-After honored
  • Exponential backoff with jitter; attempts capped
  • Circuit breaker per target
  • Conditional requests and caching enabled
  • Assets blocked where not needed
  • Browser rendering only where required
  • Proactive session rotation on a request budget
  • ASN distribution monitored
  • Realistic headers, cookies maintained within sessions