Back to blog

Guide

Proxy Scraping Workflow Example: End to End

Use this proxy scraping workflow example to plan targets, rotate IPs, manage retries, and collect compliant, dependable web data at scale safely.

This is a complete scraping workflow from planning through delivery — the sequence of decisions and the reasoning behind each, applied to a concrete scenario. The specifics are a competitor catalog monitoring job, but the structure transfers to most collection work.

The scenario: track product listings across four competitor sites in three markets (US, Germany, Brazil), roughly 12,000 SKUs per market, refreshed daily. Output feeds a merchandising dashboard.

Phase 1: Scope and Compliance

Before any code, settle what is in scope and whether it should be.

Define the data. Per listing: SKU, title, price, currency, availability, category, and collection timestamp. Nothing else. A narrow field list keeps you out of ambiguous territory and reduces page weight requirements.

Confirm public accessibility. Every target page must be reachable without authentication. Pricing behind a login, account-specific pricing, and cart-derived pricing are out of scope — not because they are technically harder, but because collecting them changes the legal character of the work.

Check robots.txt per domain. Automate the check rather than doing it once by hand, and exclude disallowed paths. Product listing pages are typically not disallowed; verify rather than assume, and re-verify periodically since these files change.

Exclude personal data. No reviews with author identifiers, no seller contact details, no user-generated content. If the parser can reach it, make the exclusion explicit in the field list rather than relying on it not being extracted.

Record the decisions. A short written scope — what, where, why, what was excluded and on what basis — is what makes the operation defensible later. It costs twenty minutes once.

Phase 2: Target Classification

Classify each target before configuring anything, because this drives every subsequent decision.

Run a manual reconnaissance pass per site: load pages normally, load them through a datacenter proxy, load them through a residential proxy, and observe what differs. Then measure per-IP tolerance by sending requests from a single IP at a steady rate until the first block signal, repeated across a dozen IPs.

For this scenario the result:

TargetTierIP type neededJS requiredMeasured tolerance
Site APermissiveDatacenterNo~400/IP/hr
Site BModerateResidentialNo~90/IP/hr
Site CModerateResidentialYes~70/IP/hr
Site DStrictResidentialYes~10/IP/hr

Site D requires more IPs than the other three combined despite similar volume. That asymmetry is typical, and discovering it now rather than in production is the entire point of this phase.

Phase 3: Proxy Configuration

Routing follows directly from the classification.

Site A — datacenter proxies, per-request rotation, 150 concurrent threads, ~250 ms spacing. Cheapest bandwidth, highest throughput.

Site B — residential, country-targeted, per-request rotation, 60 threads per market, ~1 s spacing with jitter.

Site C — residential, country-targeted, sticky sessions at 5 minutes (listing pages require a short navigation sequence), 40 threads per market, ~1.5 s spacing.

Site D — residential, country-targeted, sticky sessions, full browser with stealth configuration, 15 threads per market, randomized 4–9 s intervals. Conservative by necessity.

Session identifiers encode everything needed for later diagnosis:

user-country-BR-session-catalog_siteD_br_run20261011_t07

Country is in the session string, so cross-market contamination is structurally impossible. Run ID and thread index make gateway logs joinable against application logs.

Region targeting is worth noting here. Rather than maintaining separate configurations per country when the grouping is what matters, a region parameter — LATAM, EU, tier-1 — collapses the configuration. For this job the three markets are specific enough to name individually, but a broader version covering twenty markets would use region groups and avoid maintaining a country list.

Phase 4: Geo-Verification Gate

Every session verifies its location before collecting anything.

def open_session(country, city=None):
    for _ in range(4):
        sid = uuid.uuid4().hex[:12]
        proxies = build_proxies(country, city, sid)
        try:
            r = requests.get("https://ipcheck.flameproxies.com",
                             proxies=proxies, timeout=12)
            data = r.json()
        except Exception:
            continue
        if data.get("country_code") == country:
            if not city or city.lower() in (data.get("city") or "").lower():
                return proxies, sid, data
    raise RuntimeError(f"Could not obtain verified session for {country}/{city}")

This costs one lightweight request per session. Without it, a session that lands in the wrong country produces wrong-market prices that look entirely valid downstream — no error, no exception, just quietly incorrect data attributed to the wrong market. For a job whose entire purpose is cross-market comparison, that failure is total and invisible.

Phase 5: Fetch Layer

The fetch layer handles classification, retries, and rotation.

ROTATE_ON = {"blocked", "soft_block", "soft_block_suspect", "network_error"}
KEEP_ON   = {"target_error"}
 
def fetch(url, country, city, policy, max_attempts=3):
    proxies, sid, _ = open_session(country, city)
 
    for attempt in range(max_attempts):
        try:
            resp = requests.get(url, proxies=proxies,
                                headers=headers_for(country), timeout=30)
        except requests.exceptions.RequestException:
            resp = None
 
        status = classify(resp, required_marker=policy["marker"])
 
        if status == "ok":
            return resp, status, sid
        if status == "proxy_auth_failed":
            raise RuntimeError("Credential or parameter syntax error")
 
        if status == "blocked" and resp is not None:
            ra = resp.headers.get("Retry-After")
            if ra and ra.isdigit():
                time.sleep(min(int(ra), 60))
 
        if status in ROTATE_ON:
            proxies, sid, _ = open_session(country, city)
        elif status not in KEEP_ON:
            break
 
        time.sleep(min(30, 2 ** attempt) * random.uniform(0.5, 1.0))
 
    return None, status, sid

The rules that matter: never retry a 403 or 429 on the same IP, keep the session on a target 5xx since rotation does not fix a server error, and raise immediately on a credential failure rather than repeating it across the entire queue.

Classification treats a 200 containing a challenge page as a failure. Without that, challenge pages get stored as records and the success rate looks healthy while the dataset fills with garbage.

Phase 6: Reducing Request Count

Cheaper than any rotation strategy:

Conditional requests. Catalog pages change slowly. If-None-Match and If-Modified-Since yield 304s that cost almost nothing and do not consume per-IP budget the way a full fetch does. On this job it removes roughly half of real traffic after the first run.

Asset blocking for the browser-rendered targets (C and D). Images, fonts, media, and stylesheets are aborted. A listing page drops from ~2 MB to ~180 KB.

Volatility-weighted scheduling. SKUs whose prices have historically moved get refreshed first and most often. Stable SKUs drop to every third day. Same business value, substantially fewer requests.

Phase 7: Parse and Validate

Parse per target, then validate before anything is stored:

  • Price is a positive number within a plausible range for the category
  • Currency matches the expected currency for the market
  • Availability is one of the recognized values
  • SKU matches the expected format
  • Title is non-empty and not a generic error string

Records failing validation are flagged, not discarded. A sudden rise in validation failures is the earliest signal that a target changed its markup, and discarding the evidence destroys your ability to diagnose it.

Price anomalies — changes above 30% from the previous observation — are flagged for review. Most turn out to be genuine promotions; a minority reveal parsing errors that passed validation.

Phase 8: Monitoring

Instrumented per target and per market, not in aggregate:

  • Valid record rate — parsed, validated records over attempts. Alert on a 10-point drop against each target's own 7-day baseline.
  • Error class mix — a shift from timeouts to soft blocks is a different problem than the same total error rate suggests.
  • Soft block rate — tracked separately and alerted on a low absolute floor.
  • Bandwidth per valid record — the cost metric; a doubling means blocks, page weight changes, or parser breakage.
  • Geo-verification failure rate — rising means pool drift in that market.
  • p99 latency at flat error rate — the silent throttling signature.

A continuous canary against a neutral permissive target through each proxy type separates "proxy layer degraded" from "my targets changed" in one glance during an incident.

Failure response bodies are sampled and retained — a few hundred bytes per failure per target per hour. When Site D changes its challenge mechanism, that sample is what tells you what the new one looks like.

Phase 9: Output

Daily export, partitioned by market and target, each record carrying:

{
  "sku": "...",
  "site": "site_c",
  "market": "DE",
  "verified_city": "Frankfurt",
  "price": 49.99,
  "currency": "EUR",
  "availability": "in_stock",
  "collected_at": "2026-10-11T06:20:00Z",
  "validation": "ok"
}

verified_city is the location actually confirmed at session start, not the one requested. Months later, this is what distinguishes a genuine market price change from a session that quietly drifted to the wrong location.

A quality report accompanies each export: valid rate per target and market, flagged record counts, anomalies, guardrails that fired, and known issues.

Why Throughput Shapes the Whole Design

Notice the pattern across Phase 3: every target uses many concurrent sessions, each sending requests slowly. Site D runs 15 threads per market at 4–9 second intervals — gentle per IP, parallel across many.

That shape is deliberate, and it is the only way to collect from a strict target at volume without triggering throttling. The alternative — fewer sessions each pushing harder — hits per-IP tolerance immediately.

It is also the shape a concurrency cap makes impossible. With 45 threads across three markets for Site D alone, plus 120 for Site C, 180 for Site B, and 150 for Site A, this single job wants roughly 500 concurrent sessions. On a provider metering concurrency by plan tier, you either buy a tier you do not otherwise need or compress the sessions and get blocked.

FlameProxies applies no concurrency cap and no request-rate limits, so the wide-and-gentle configuration above is sized by the job's requirements rather than by an allowance — and the geo-verification gate on every session runs alongside collection rather than competing with it for throughput.