Guide
Web Scraping Proxy Example in Python (requests)
See a web scraping proxy example in Python. Use rotation and country targeting to reduce blocks, control retries, and protect data collection at scale.

This is a working example of proxy-based scraping in Python with requests — rotation, country targeting, retry handling that distinguishes failure types, and the response validation that keeps challenge pages out of your dataset. It builds up from the minimal case to something you would actually run.
The Minimal Case
A proxy in requests is a dict mapping scheme to proxy URL:
import requests
PROXY_HOST = "gateway.provider.com"
PROXY_PORT = 8080
USERNAME = "user12345"
PASSWORD = "your_password"
proxy_url = f"http://{USERNAME}:{PASSWORD}@{PROXY_HOST}:{PROXY_PORT}"
proxies = {"http": proxy_url, "https": proxy_url}
resp = requests.get("https://api.ipify.org?format=json", proxies=proxies, timeout=15)
print(resp.json())Note that the https key takes an http:// proxy URL. That is correct and routinely confusing — the scheme in the key is the scheme of the destination, while the scheme in the value is how you reach the proxy. HTTPS destinations are handled through an HTTP CONNECT tunnel.
Adding Country Targeting and Sessions
Residential gateways typically accept targeting parameters encoded in the username. The exact syntax varies by provider; this shape is representative:
def build_proxies(country=None, city=None, session_id=None):
parts = [USERNAME]
if country:
parts.append(f"country-{country}")
if city:
parts.append(f"city-{city}")
if session_id:
parts.append(f"session-{session_id}")
user = "-".join(parts)
url = f"http://{user}:{PASSWORD}@{PROXY_HOST}:{PROXY_PORT}"
return {"http": url, "https": url}Omitting session_id gives per-request rotation. Supplying a stable one gives a sticky session on the same IP. Two requests with the same session ID should return the same IP; different IDs should return different IPs — worth asserting once during setup.
Verifying Geography Before Collecting
For location-sensitive work, confirm the IP landed where you asked before you collect anything through it:
def verify_location(proxies, want_country, want_city=None, timeout=15):
try:
r = requests.get("https://ipapi.co/json/", proxies=proxies, timeout=timeout)
data = r.json()
except Exception:
return False
if data.get("country_code") != want_country:
return False
if want_city and want_city.lower() not in (data.get("city") or "").lower():
return False
return TrueA session that fails this check should be discarded and re-requested, not used. Wrong-market data produces no error anywhere downstream — it just quietly misrepresents the market you thought you were measuring.
Classifying Responses
Status code alone is not enough. A 200 containing a captcha is a failure, and treating it as success is how challenge pages end up stored as records:
CHALLENGE_MARKERS = (
"captcha", "unusual traffic", "verify you are human",
"access denied", "cf-challenge", "are you a robot",
)
def classify(resp, min_length=2000, required_marker=None):
if resp is None:
return "network_error"
if resp.status_code == 407:
return "proxy_auth_failed"
if resp.status_code in (403, 429):
return "blocked"
if resp.status_code >= 500:
return "target_error"
if resp.status_code != 200:
return "unexpected_status"
body = resp.text
low = body.lower()
if any(m in low for m in CHALLENGE_MARKERS):
return "soft_block"
if len(body) < min_length:
return "soft_block_suspect"
if required_marker and required_marker not in body:
return "content_anomaly"
return "ok"required_marker is a string you expect on every valid page for a given target — a container class, a label, anything stable. It catches page restructures that would otherwise silently break parsing.
Retry Logic That Responds to the Failure Type
The important rule: rotate the IP on a block, keep the IP on a target server error, and never retry a 429 on the same IP.
import random
import time
import uuid
def backoff(attempt, base=1.0, cap=30.0):
return min(cap, base * (2 ** attempt)) * random.uniform(0.5, 1.0)
ROTATE_ON = {"blocked", "soft_block", "soft_block_suspect", "network_error"}
KEEP_SESSION_ON = {"target_error"}
def fetch(url, country, city=None, required_marker=None,
max_attempts=3, headers=None):
session_id = uuid.uuid4().hex[:12]
proxies = build_proxies(country, city, session_id)
for attempt in range(max_attempts):
try:
resp = requests.get(url, proxies=proxies, headers=headers, timeout=30)
except requests.exceptions.RequestException:
resp = None
status = classify(resp, required_marker=required_marker)
if status == "ok":
return resp, status
if status == "proxy_auth_failed":
raise RuntimeError("Proxy credentials or parameter syntax rejected")
if status == "blocked" and resp is not None:
retry_after = resp.headers.get("Retry-After")
if retry_after and retry_after.isdigit():
time.sleep(min(int(retry_after), 60))
if status in ROTATE_ON:
session_id = uuid.uuid4().hex[:12]
proxies = build_proxies(country, city, session_id)
elif status not in KEEP_SESSION_ON:
break
time.sleep(backoff(attempt))
return None, statusRaising immediately on proxy_auth_failed matters — retrying a credential problem across every URL in a queue produces thousands of identical failures and obscures the actual cause.
Realistic Headers
Default requests headers are an obvious automation signature. Send something plausible:
DEFAULT_HEADERS = {
"User-Agent": (
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/141.0.0.0 Safari/537.36"
),
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Encoding": "gzip, deflate, br",
"Connection": "keep-alive",
"Upgrade-Insecure-Requests": "1",
}
LOCALE_BY_COUNTRY = {
"US": "en-US,en;q=0.9",
"GB": "en-GB,en;q=0.9",
"DE": "de-DE,de;q=0.9,en;q=0.8",
"FR": "fr-FR,fr;q=0.9,en;q=0.8",
}
def headers_for(country):
h = dict(DEFAULT_HEADERS)
h["Accept-Language"] = LOCALE_BY_COUNTRY.get(country, "en-US,en;q=0.9")
return hMatching Accept-Language to the proxy's country is not cosmetic. A German IP sending en-US is an incoherent profile, and some targets resolve that incoherence in ways that change what you receive.
Pacing and Concurrency
Wide parallelism with gentle per-IP pacing is the pattern that avoids throttling. A thread pool with a per-target rate limit:
from concurrent.futures import ThreadPoolExecutor, as_completed
from threading import Semaphore
CONCURRENCY = 50
PER_REQUEST_DELAY = 1.5 # seconds, per worker
JITTER = 0.4
limiter = Semaphore(CONCURRENCY)
def worker(url, country, city=None, required_marker=None):
with limiter:
time.sleep(PER_REQUEST_DELAY * (1 + random.uniform(-JITTER, JITTER)))
return url, fetch(
url, country, city,
required_marker=required_marker,
headers=headers_for(country),
)
def run(urls, country, city=None, required_marker=None):
results, failures = {}, {}
with ThreadPoolExecutor(max_workers=CONCURRENCY) as pool:
futures = [
pool.submit(worker, u, country, city, required_marker)
for u in urls
]
for fut in as_completed(futures):
url, (resp, status) = fut.result()
if status == "ok":
results[url] = resp.text
else:
failures[url] = status
return results, failuresEach worker holds its own session and sends paced requests, so no single IP accumulates much history. Raising CONCURRENCY scales throughput by adding IPs rather than by pushing any one harder — which is the correct direction, and the reason a provider's concurrency cap becomes your ceiling if it has one.
Reducing Request Count
Cheaper than any rotation strategy: do not send the request.
def fetch_conditional(url, etag=None, last_modified=None, **kwargs):
headers = kwargs.pop("headers", {}) or {}
if etag:
headers["If-None-Match"] = etag
if last_modified:
headers["If-Modified-Since"] = last_modified
resp, status = fetch(url, headers=headers, **kwargs)
if resp is not None and resp.status_code == 304:
return None, "not_modified"
return resp, statusA 304 costs almost no bandwidth and does not consume per-IP budget the way a full fetch does. On a catalog that changes slowly, conditional requests can cut real traffic by most of its volume.
Logging for Diagnosis
Record the dimensions you will want when a success rate drops:
import json
from datetime import datetime, timezone
def log_attempt(url, country, city, session_id, status, elapsed, body=None):
record = {
"ts": datetime.now(timezone.utc).isoformat(),
"url": url,
"country": country,
"city": city,
"session_id": session_id,
"status": status,
"elapsed_s": round(elapsed, 3),
}
if status != "ok" and body:
record["body_prefix"] = body[:500]
print(json.dumps(record))Keeping a truncated body on failures is what later lets you distinguish a captcha from a redirect from a restructured page. Without it you are guessing from status codes.
Putting It Together
urls = [f"https://example-target.com/product/{i}" for i in range(1, 501)]
results, failures = run(
urls,
country="DE",
city="Munich",
required_marker='class="product-detail"',
)
print(f"ok: {len(results)} failed: {len(failures)}")
from collections import Counter
print(Counter(failures.values()))The failure breakdown by class is the output worth reading. A pile of blocked means pacing or IP quality. A pile of content_anomaly means the target changed its markup and your parser needs attention — a completely different problem with a completely different fix.
FlameProxies works with this pattern directly: HTTP/HTTPS and SOCKS5 endpoints, username-encoded country and city targeting, sticky and rotating sessions, and no concurrency cap or request-rate limit — so CONCURRENCY above is bounded by your own machine rather than by a thread allowance, which is what makes the wide-and-gentle approach viable at volume.