Guide
AI Data Acquisition: Proxies for Training and Eval Corpora
AI data acquisition proxies help teams collect geo-specific web data at scale while controlling retries, cost, compliance, and IP rotation across markets.

Building training corpora and evaluation sets from the web is a different problem than operational scraping, and the differences change the infrastructure requirements. Operational scraping collects a known schema from known targets on a schedule. Data acquisition for AI collects heterogeneous content from a long tail of sources, once or in large irregular batches, at volumes where cost per gigabyte stops being an abstraction. It also carries provenance requirements that operational collection usually does not.
Proxies sit underneath all of it. This covers what changes at that scale and how to configure for it.
What Makes This Different
Volume is categorically larger. A price monitoring job moves hundreds of gigabytes monthly. A corpus build can move hundreds of terabytes. At that scale, small inefficiencies compound into large numbers — a 15% retry rate is not a quality annoyance, it is a line item.
The target list is long and heterogeneous. Instead of five known sites, thousands of domains with wildly varying structure, page weight, protection level, and content quality. You cannot hand-classify each one.
Collection is bursty, not steady. A corpus build runs hard for days or weeks, then stops. Infrastructure sized for the average is wrong; what matters is peak parallelism during the build window.
Geographic diversity is a data property, not a logistics detail. If your corpus is collected entirely from one region's IPs, it is systematically skewed toward what that region is served — a real representativeness problem, not just a coverage gap.
Provenance has to be recorded at collection time. Where each document came from, when, under what access conditions, and what the source's terms were. Reconstructing this later ranges from expensive to impossible.
Geographic Diversity as Data Quality
This is the point most often missed, so it is worth being concrete.
Collect a multilingual corpus entirely through US IPs and several things happen. Sites serve English versions rather than local-language versions. Regional content gets geo-filtered out. Localized pricing, availability, and editorial content reflect the US market. Some sites block or degrade foreign traffic entirely, so their content is underrepresented relative to its actual prevalence.
The resulting corpus is not a sample of the web — it is a sample of what the web shows an American visitor. For a model intended to work across markets, that skew is baked into the training data and is very hard to detect afterward.
Collecting through IPs in the regions whose content you want produces the content those regions actually see. For multilingual or market-specific corpora, this is a correctness requirement.
Region targeting helps operationally here. Rather than maintaining and rotating a country list, region groups — EU, LATAM, tier-1 markets — let you express coverage goals directly and distribute collection across the regions you need without per-country bookkeeping.
Tiering a Long Target List
You cannot classify thousands of domains manually. Automate the triage.
Run a cheap probe against each domain before full collection: one request through a datacenter proxy, one through residential, comparing status codes, response sizes, and whether a challenge page appears. From that, bucket automatically:
Tier 1 — open. Returns content to datacenter IPs with no challenge. The large majority of a typical long-tail list. Route to datacenter proxies, no rendering, high concurrency. Cheapest bandwidth by a wide margin.
Tier 2 — IP-filtered. Blocks datacenter, serves residential. Route to residential, country or region targeted, moderate concurrency.
Tier 3 — protected. Challenges both, or requires JavaScript to render content. Route to residential with full browser rendering, low per-IP rates.
Tier 4 — not worth it. Hard-blocked, or cost per usable document is implausible. Record the decision and skip. Being willing to drop sources is a cost control; chasing the last few percent of a long tail can consume more budget than the bulk of the corpus.
Re-run triage periodically. Sites change tiers, and a domain stuck in Tier 3 that has since loosened is pure overspend.
The economics of this are stark. If 80% of your list is Tier 1 and you route everything through residential because it is simpler, you are paying multiples for the large majority of your volume with no benefit.
Cost Control at Corpus Scale
Strip bytes you will not use. Most corpus work needs text, not assets. For browser-rendered targets, abort images, fonts, media, and stylesheets. For plain HTTP, request with compression and avoid rendering entirely where the HTML carries the content. The difference between fetching a full page and fetching its HTML is frequently an order of magnitude.
Deduplicate before fetching, not after. URL-level dedup is obvious. Beyond that, many corpus builds fetch near-identical pages — paginated variants, print versions, tracking-parameter duplicates of the same document. Normalize URLs and apply content hashing early. Documents discarded at the dedup stage were paid for in full.
Sample before committing to a source. For a large domain, fetch a few hundred documents and assess quality and uniqueness before crawling the whole thing. Many large sites contribute far less unique value than their page count suggests.
Cap per-domain spend. A budget ceiling per domain prevents one pathological source — infinite pagination, calendar pages, generated URL space — from consuming a disproportionate share. This failure is common and silent without a cap.
Track bytes per retained document. Not bytes per request, not bytes per fetch — per document that survives dedup and quality filtering. This is the number that maps to corpus cost, and it exposes sources that are expensive relative to what they contribute.
Rotation and Pacing for Heterogeneous Targets
With thousands of domains you cannot tune per-IP tolerance individually. Use adaptive defaults.
Start every unknown domain conservatively — modest concurrency, spaced requests with jitter. Track per-domain success rate. Increase concurrency for domains sustaining a high success rate; back off automatically on domains showing blocks. A simple feedback loop handles the long tail far better than static configuration.
Per-domain circuit breakers matter more here than in operational scraping, because nobody is watching individual domains. A domain that starts blocking should be cut off automatically rather than consuming retry bandwidth for hours.
Spread load across ASNs, not just addresses. Many targets aggregate reputation at network level, so ASN diversity does more for sustained success than raw IP count. ASN targeting gives you direct control here — steering distribution rather than hoping rotation spreads it adequately.
Provenance and Compliance
Record at collection time, because reconstruction is not feasible:
- Source URL and final URL after redirects
- Collection timestamp
- HTTP status and content type
- Whether the content was publicly accessible without authentication
- The robots.txt directive state for that path at collection time
- Content hash for dedup and later deduplication audits
- The collecting job and configuration version
Operating principles worth holding to:
Public content only. No authentication, no paywall circumvention, no access to content the site does not serve to ordinary visitors.
Respect robots.txt. Check it per path, automatically, and record the result. For corpus work specifically, robots directives are frequently the site's stated position on automated collection, and ignoring them is both a legal and reputational exposure.
Exclude personal data by policy. Build exclusion into the collection layer rather than relying on downstream filtering.
Rate limit as a courtesy, not only to avoid blocks. Corpus builds can generate load far beyond what a small site expects. Conservative pacing on small domains is the right default independent of whether they would block you.
Source your proxies ethically. Residential networks built on IPs whose owners did not knowingly consent create legal exposure that transfers to you — and for AI training data, where provenance is increasingly scrutinized, that exposure attaches to the dataset itself. Consent-based sourcing is a dataset property, not just a vendor detail.
Evaluation Sets Have Different Requirements
Worth separating from training corpora.
Eval sets are smaller but need higher precision and often tighter geographic control. If you are evaluating model behavior on localized content, the eval data must genuinely come from those locales — collected through IPs in the right markets, with coherent language and timezone signals so the content served matches what a local user sees.
Eval sets also need reproducibility. Record enough configuration detail that a later collection run can be compared meaningfully against the original. The verified geographic location of each collection session belongs in the record, not just the requested one.
Throughput During the Build Window
Corpus collection is defined by its burst. A build that wants to finish in two weeks rather than four months needs very wide parallelism during that window — tens of thousands of concurrent sessions is not unusual at the upper end.
This is where provider architecture becomes the binding constraint. Everything above points toward many concurrent sessions each paced gently, which is both the way to avoid blocks and the way to finish in reasonable time. A concurrency cap forces the opposite: fewer sessions pushing harder, which raises block rates on exactly the protected sources that are hardest to collect, and stretches the build window regardless.
It also means instrumentation competes with collection. Geo-verification, triage probes, and canary monitoring all consume throughput. On a metered provider, teams cut instrumentation to preserve collection capacity — and then cannot explain their data quality problems.
FlameProxies applies no concurrency cap and no request-rate limits, with users sustaining over 300,000 requests per second in production. For corpus work that means the build window is sized by your own infrastructure rather than by a plan tier, and verification runs alongside collection instead of against it. Combined with 80M+ ethically sourced residential IPs across 180+ countries, region and ASN targeting, and non-expiring bandwidth at $0.50/GB, the geographic diversity that makes a corpus representative is available without the throughput trade-off.
Summary Checklist
- Automated tier triage before full collection, re-run periodically
- Datacenter for Tier 1; reserve residential for targets that need it
- Geographic distribution matched to corpus coverage goals
- Region targeting for broad coverage; ASN targeting for distribution control
- Assets stripped; rendering only where required
- Dedup before fetch; content hashing early
- Per-domain spend caps and circuit breakers
- Adaptive pacing with per-domain feedback
- Bytes per retained document tracked per source
- Provenance recorded at collection time
- robots.txt checked per path and logged
- Verified geographic location stored per session
- Ethically sourced residential IPs documented for dataset provenance