Geo Experiment / Matched-Market Incrementality Test
When you need this method
You shift budget between channels based on attribution numbers derived from observational data. Buyers who would have converted anyway are credited to the last touchpoint, so whichever channel sits closest to the close looks best. You therefore do not know which part of your demand is genuinely incremental and which is merely reassigned.
Approach
- 1Partition the target market into non-overlapping geos, large enough for reliable geo-targeted delivery, small enough to yield a sufficient number of units.
- 2Fix the response metric, for example inquiries, qualified opportunities or revenue, and assemble pretest-period data for each geo.
- 3Rank the geos by their pretest level, cut the ranked list into groups of size M, and draw one geo per group at random into the treatment group; the test fraction is then 1/M.
- 4Compute the achievable precision before starting: build pseudo pretest and test periods from the historical data, estimate the variance of the effect coefficient, and average it across many random assignments. If the half-width exceeds the effect you expect, this design cannot answer the question.
- 5Run the test period and change the budget only in the treatment group; campaign structure, bids and creatives stay constant in both groups.
- 6Analyse with the weighted regression model that regresses the test-period response on the pretest response and the spend differential; the coefficient on the spend differential is the incremental effect per euro spent.
- 7Extend the test period beyond the end of the budget change until the lagged effect has worked through; with offline and pipeline metrics this takes time.
Typical application
A typical B2B SaaS in the HR space wants to cut its paid search budget because the attribution model credits the channel with little. Instead of cutting, the team divides the German market into postcode clusters, ranks them by qualified inquiries in the pretest period, and draws one cluster per group of three into the treatment group. The precision calculation shows that a one-fifth reduction would disappear into the noise, so the intervention is made larger and the test period is stretched to eight weeks. The measured drop in qualified inquiries turns out to be considerably larger than what the attribution model had assigned to the channel. The budget stays, and the internal debate moves from which attribution model is right to which intervention gets tested next.
Limits and counter-indications
The method presupposes that advertising can be targeted geographically and that both spend and the response metric can be attributed cleanly per region; channels without geo control are out. It needs enough regions with enough volume, because with few and highly heterogeneous units the estimate becomes unreliable, which is why Chen and Au propose a more robust estimator. In many B2B cases the precision calculation already fails, since the achievable confidence width exceeds the expected effect; that is a solid finding, but it means the test does not happen. With long buying cycles the test period must cover the lag, and the primary source itself notes that the precision forecast becomes overly optimistic once the same pretest data are reused for long test periods. The result holds for the spend range, the period and the region tested, it is not a general elasticity and says nothing about other channels.
How to measure impact
Track the incremental effect per euro spent together with the half-width of its confidence interval, and set it against the attributed effect for the same period. The gap between the two numbers measures how far your previous budget decisions were skewed.
Related methods
Sources
- 1.Vaver & Koehler: Measuring Ad Effectiveness Using Geo Experiments, Google Inc., 2011 (opens in a new tab) · Google Inc. (Google Research) · 2011 · investment, consulting and analyst firms, industry bodies and public agencies · describes the methodCarries the full procedure: non-overlapping geos, random assignment to test and control, the weighted linear model with weights of 1/pretest value, ROAS as the coefficient on the spend differential, the variance formula, and the up-front estimate of confidence width from pseudo periods. It states explicitly that grouping geos by size before assignment narrows the confidence interval by roughly ten percent or more, and that the test period must run past the end of the budget change because offline sales respond with a lag. Limit: a Google research report without peer review, with no industry, sample size or transferability to B2B stated; the authors themselves note that no measurement method works in every situation.
- 2.Gordon, Zettelmeyer, Bhargava & Chapsky: A Comparison of Approaches to Advertising Measurement, Marketing Science 38(2), 2019 (opens in a new tab) · Kellogg School of Management / INFORMS Marketing Science · 2019 · academic and scholarly literature · supports the underlying mechanismEstablishes why an experiment is needed at all. The authors compare fifteen randomized advertising experiments on Facebook, covering some 500 million user-experiment observations, against common observational models run on the same data. The observational models mostly overstate the effect and in some cases badly understate it; in half of the studies the estimated lift in purchase outcomes is off by a factor of three across all methods. Limit: this was measured on one platform with unusually rich user-level data, not at geo level; the study evidences the failure of observational methods, not the quality of the geo experiment.
- 3.Kerman, Wang & Vaver: Estimating Ad Effectiveness using Geo Experiments in a Time-Based Regression Framework, Google Inc., 2017 (opens in a new tab) · Google Inc. (Google Research) · 2017 · investment, consulting and analyst firms, industry bodies and public agencies · describes the methodCarries the variant for small numbers of units. The authors state that geo-based regression is barely applicable, or not applicable at all, when only few geographic units exist, and set alongside it a time-based regression that predicts the time series of the counterfactual market response and estimates the cumulative causal effect from it. Limit: again a Google report without peer review; moving the burden from the number of regions to the quality of a time-series forecast is a trade, not the removal of a requirement.
- 4.Chen & Au: Robust causal inference for incremental return on ad spend with randomized paired geo experiments, Annals of Applied Statistics 16(1), 2022 (opens in a new tab) · Institute of Mathematical Statistics / Annals of Applied Statistics · 2022 · academic and scholarly literature · limits the methodNames the weaknesses of the standard procedure in a peer-reviewed venue: with a small number of highly heterogeneous regions, reliable estimates of the incremental effect are hard to obtain, budget constraints create interference between units, and the usual distributional assumptions are violated. The authors propose Trimmed Match, a distribution-free and more robust estimator. Limit: the paper replaces the analysis formula, not the method; the demands on design and data remain.
Origin: Vaver & Koehler (Google Research)