Guide
How to Use Proxies for Ad Testing Across Markets and Devices
Learn how to use proxies for ad testing across markets, devices, and placements, with cleaner verification data and fewer location-based blind spots daily.

Ad testing and ad verification overlap but are not the same discipline. Verification asks whether a live campaign delivered as booked. Testing asks what a campaign looks like under conditions you choose — different markets, devices, placements, and audience profiles — often before or during a campaign rather than as a compliance record afterward. Proxies are what make that range of conditions reachable, because nearly every variable that changes what an ad platform serves you is tied to the IP the request comes from.
This guide covers how to structure proxy-based ad testing so the results are clean enough to act on.
What Changes What You See
Before configuring anything, it helps to be precise about which signals drive ad delivery differences, because your test design has to control each one.
IP geolocation determines geographic targeting eligibility, which currency and pricing appear in dynamic creatives, and which regional campaigns are in the auction at all.
IP type classification determines whether you are treated as a real consumer or as infrastructure. Datacenter IPs frequently receive different inventory, reduced fill, or no ads at all on geo-targeted campaigns. This is the single most common reason ad testing produces results that do not match reality.
Device and user agent determine which creative formats are eligible — mobile interstitials, responsive display sizes, app-install units. A desktop user agent will never surface mobile-only inventory regardless of your proxy configuration.
Session history and cookies drive frequency capping, retargeting eligibility, and sequential creative delivery. A clean session sees the first-touch experience; a session with history sees something else entirely.
Language and timezone headers cross-reference against IP geolocation. Sophisticated platforms treat a mismatch — a German IP with en-US language and a US timezone — as a signal that the visitor is not a genuine local user, and may adjust delivery accordingly.
A test that varies one of these while accidentally varying another produces results you cannot attribute. Control all five explicitly.
Structuring the Test Matrix
Ad testing is a matrix problem. The dimensions that typically matter:
Markets: the countries, and where relevant the cities, whose ad experience you need to observe.
Devices: desktop, mobile web, tablet — each surfacing different inventory and creative formats.
Placements: the specific publisher pages, search results, feeds, or app contexts where the ad should appear.
Session profiles: clean first-visit sessions versus sessions with prior browsing history, for testing retargeting and frequency behavior.
Each cell of the matrix is a separate proxy session with its own configuration. A test covering 5 markets × 3 devices × 4 placements × 2 session profiles is 120 cells — and each should be run multiple times, because ad delivery is probabilistic. A single observation per cell tells you almost nothing.
That multiplication is why throughput matters in ad testing more than people expect. A thorough test matrix run across meaningful repetitions produces tens of thousands of page loads, each requiring full JavaScript rendering and ad auction resolution.
Proxy Configuration Per Test Cell
Market dimension
Residential IPs with targeting matched to the campaign's own targeting granularity. If the campaign targets a metro, test at the metro level; country-level targeting will not reproduce the audience's experience for a city-targeted campaign.
Gate every session on geo-verification before it records anything. A session that fails the check has to be discarded rather than logged, or you contaminate the market dimension of your matrix.
Device dimension
The device dimension is largely browser configuration rather than proxy configuration, but the two must be consistent. A mobile test should pair a mobile user agent and viewport with a residential IP — and if you are testing mobile carrier inventory specifically, a mobile proxy rather than a broadband residential IP, since carrier ASN is itself a targeting signal on some platforms.
Set viewport, device pixel ratio, touch support, and user agent together. Platforms cross-check these, and a desktop viewport with a mobile user agent is a recognizable inconsistency.
Session profile dimension
For clean first-touch testing, start each session with cleared cookies and storage, and use a fresh IP that your operation has not recently used against that platform. Reusing an IP that already has frequency history against the campaign gives you a second-impression experience labeled as a first.
For retargeting and sequencing tests, the opposite: hold a sticky session, build the intended history deliberately, and keep the same IP throughout. A mid-test IP change breaks the continuity the test depends on.
Consistency across the session
Language headers and timezone should match the market being tested. Accept-Language: de-DE with a German IP and a Europe/Berlin timezone is coherent; mixing them is a detectable inconsistency that can shift what you are served and invalidate the cell.
Reading the Results Correctly
Ad delivery is probabilistic, and this is the most common source of wrong conclusions in ad testing.
The same user in the same location loading the same page twice can see different ads, because of auction dynamics, pacing, budget exhaustion, frequency caps, and live A/B tests on the platform side. A single observation showing your ad absent does not mean the ad is not delivering. It means it did not win that auction.
Design around this:
Repeat every cell. Ten to twenty observations per cell is a reasonable floor for drawing conclusions about presence and share of voice. One is noise.
Report rates, not binaries. "Ad appeared in 17 of 20 checks" is a finding. "Ad appeared" is an anecdote.
Spread observations across time. Campaign pacing means delivery varies by hour. A test run entirely in one 20-minute window measures that window, not the campaign.
Separate absence from blocking. An ad slot that renders empty is a delivery outcome. A page that returned a captcha, timed out, or failed to execute ad JavaScript is a collection failure and must be excluded rather than counted as non-delivery. Conflating the two inflates your reported non-delivery rate with your own infrastructure's failures.
That last point is where proxy quality translates directly into data quality. Throttled or blocked requests that half-render produce phantom findings — the test reports that a campaign is not serving in a market when the real cause was your own collection layer stalling.
Common Mistakes
Testing through datacenter IPs. The fastest way to produce ad testing data that does not reflect reality. Many platforms serve reduced or no inventory to datacenter ranges. Use residential IPs and confirm the classification.
Country targeting for city-targeted campaigns. Produces the national-average experience rather than the local one, and will systematically under-report delivery for metro-targeted campaigns.
Reusing sessions across matrix cells. Cookie and frequency history from one cell leaks into the next, and the second cell's results are no longer clean.
Ignoring the JavaScript requirement. Most ad creatives render through JavaScript and complete an auction after initial page load. Plain HTTP requests see an empty slot every time. Full browser automation with an adequate post-load wait is mandatory.
Under-sampling. By far the most frequent analytical error. Probabilistic delivery demands repetition, and conclusions drawn from single observations per cell are unreliable regardless of how well the proxy layer is configured.
Scaling the Matrix
A modest test matrix run properly generates substantial traffic. 120 cells × 15 repetitions is 1,800 full browser sessions, each loading a real publisher page with complete ad rendering — on the order of 20,000 to 30,000 HTTP requests, concentrated into whatever window the test needs to run in.
Run that matrix against a provider that caps concurrency and the test either stretches across a window long enough that campaign conditions change underneath it, or gets throttled into producing the phantom-absence findings described above. The constraint is not bandwidth; it is how many browser sessions you can drive in parallel.
FlameProxies places no limits on concurrent sessions and applies no request-rate caps, so a test matrix is sized by what your own browser infrastructure can drive rather than by a provider's throughput tier. Combined with residential IPs across 180+ countries and city-level targeting, that means the matrix you design is the matrix you can actually run — in a tight enough window that every cell reflects the same campaign conditions.