Choose the comparison design that best predicts what the treatment market would have done without the media change. That may be a real matched market, a weighted synthetic control, or—when data and test capacity allow—both.

The question every geo test must answer

A geo holdout changes advertising exposure in part of the country: pause a channel in selected markets, introduce a new strategy, or test a different spend level. The intervention alone does not prove lift. You still need a credible answer to: what would have happened in those markets if the media had not changed?

That unobserved alternative is the counterfactual. Matched-market and synthetic-control designs are two ways to construct it. Both can be useful; neither is automatically a better answer for every brand, channel, or geography.

Matched markets: compare real places

A matched-market design pairs treatment geography with one or more real markets that historically behaved similarly. Its appeal is intuitive: if Seattle has reliably tracked the selected treatment markets, Seattle can provide a tangible real-world comparison while the treatment condition changes.

But “similar on average” is not enough. Markets can look alike across a long period while diverging during the spikes and valleys that matter to a test. A weather event, local promotion, inventory issue, competitor activity, or different demand pattern can break an otherwise attractive pair. Real places preserve the messiness of the business, which is both their value and their limitation.

Synthetic control: construct a weighted comparison

A synthetic control combines several geographies into a weighted benchmark intended to resemble the treatment market’s prior behavior. It can create a very close historical fit, including when no single real market is an adequate match.

The tradeoff is capacity. A synthetic control may draw on a broad share of the country, leaving fewer untouched markets for another concurrent test. A lower-volume business may already need a large holdout to detect a meaningful result, making multiple comparisons impractical. That is a feasibility finding—not a reason to force a test.

Use historical data to validate the design

Start with candidate treatment markets, then move back through the historical data. Ask whether each comparison is appropriately sized, tracks the treatment pattern after scaling, and remains close most of the time—not merely at the average. Mean absolute percentage error (MAPE) is one useful way to inspect how far a comparison drifts from the line. It helps expose a control that averages well but is too volatile to predict the treatment market reliably.

Where volume and test capacity allow, pre-validate several comparisons: a matched-market view, a synthetic-control view, and the broader national trend. A result that is consistent across these perspectives is harder to dismiss. If the readings differ, the difference itself can show whether the test is sensitive to a fragile control choice.

Validation also includes practical execution. Define which geographies receive or lose media, check for channel bleed, choose the outcome and observation window based on the customer consideration cycle, and establish the decision rule before seeing the result. Sales and revenue tests often need roughly four to six weeks in market; longer consideration cycles and upper-funnel interventions can require a longer window.

How this complements MMM

MMM can identify a portfolio-level hypothesis: a channel may be contributing less or more incrementally than last-click reporting suggests. A geo test can check a specific causal question under defined conditions. The resulting lift evidence can calibrate the model, while the model helps identify the next question worth testing. Neither method removes the need to understand the other.

FAQs

Should we always use synthetic control?

No. It can provide an excellent historical fit, but a real matched market may be more stable or more practical. Test capacity and historical back-testing should decide.

Can a small DTC brand run a geo test?

Possibly, but it may require a large holdout or a longer observation period. If the design cannot detect a decision-relevant change credibly, a different measurement step is better.

Should a geo test use more than one comparison?

When data and capacity permit, yes. Agreement across a matched-market comparison, synthetic control, and national trend can increase confidence in the learning.

Useful measurement is a connected system: MTA vs. MMM vs. incrementality testing can help you choose the right measurement lens for the decision in front of you., why MTA and MMM disagree can help you turn conflicting readings into a testable story., whether you need MMM yet can help you assess whether the portfolio question is timely., why multiple MMM perspectives help can help you pressure-test a consequential model recommendation., and why marketing numbers disagree can help you reconcile the inputs before interpreting the result.