Dark Wave Marketing Science Ask Carlos

Measurement

Matched markets or synthetic control

A geo holdout is the cleanest causal read most advertisers can get without building anything exotic. Everything difficult about it sits in how you build the comparison.

Dark Wave Marketing Science

The mechanism is simple enough to explain in a sentence. You change media spend in one set of geographic markets, leave another set alone, and compare what happens to the business outcome in each. The difference is what the media caused.

Everything difficult about it sits in the word compare.

The markets you turned spend off in are not identical to the markets you left alone. They have different seasonality, different competitors, different weather, different local economies. If you simply average the treated markets against the untreated ones, you are attributing all of that difference to media.

The standard answer: matched markets

Most geo testing in practice is matched-market testing. You pick treatment markets, then pick control markets that resemble them on a handful of observable characteristics, usually population, historical revenue, and category behaviour. You run the test and compare the two groups.

It is cheap, it is easy to explain, and when the markets genuinely are similar it works. It is also the right tool when you only have a handful of markets to work with, because the alternative needs a larger pool to draw from.

Where it strains is in the matching itself. You are choosing controls on a few characteristics you happened to measure, at one point in time, and then assuming that similarity holds through the test. The result is sensitive to which markets you picked, and the picking is usually done with more judgment than method. Two competent analysts can select different control sets from the same market list and get different answers, and neither has a principled way to say which is right.

The stronger answer: synthetic control

Synthetic control removes the choosing. Rather than nominating a twin market and hoping, you construct a weighted composite from the whole pool of untreated markets, with the weights fitted so the composite reproduces the treated markets' actual behaviour during the pre-period, before anything changed. Any market can contribute, and most contribute nothing. The weights come out of the data instead of out of a meeting.

Two things follow from that, and both matter to a sceptical reader. The comparison group is built to match on the outcome you care about over time, not on demographics that correlate with it. And the quality of the match is visible: you can plot the composite against the treated markets for the pre-period and show how closely they tracked. The evidence for the method sits on the same chart as the result.

It also supports a check matched markets cannot easily offer. You can run the same procedure on markets that were never treated, where the true effect is known to be zero, and see how often the method finds an effect anyway. If your treated result sits inside that distribution of false positives, you do not have a finding.

The pre-period fit is the load-bearing part either way. A control that tracks well before treatment and diverges after gives you something defensible. One that never tracked well in the first place gives you a number with no claim on reality, and it will look exactly the same in a slide.

The first question to ask about any geo test result is not what the lift was. It is how well the control tracked the treatment group before the test started.

What makes a business testable

Geo design needs geographic variation to work with. The requirements are structural rather than arbitrary. You need enough distinct markets, because a single-metro advertiser has no pool to build a control from and no amount of statistics fixes that. You need enough spend in the treated markets to move something, or the effect will be smaller than the week-to-week noise. You need an outcome you can observe by geography, which online revenue with a billing postcode gives you and a national wholesale number does not.

You also need reasonable separation between markets. Media that spills across market boundaries contaminates the control group and biases the result toward zero, which is the most dangerous direction to be wrong in, because it looks like a finding. This is a live problem with broadcast overlap, with connected TV targeting that respects market boundaries loosely, and with any channel where the geography you can buy differs in shape from the geography you can measure.

If those conditions do not hold, a geo holdout is the wrong instrument. That is worth establishing early, because it is far cheaper than finding out after a quarter of testing, and because other designs exist with different strengths and a different set of things they cannot see.

Designing the comparison is only half of it. The other half is whether the result will survive contact with someone who does not want to believe it. That is the subject of the next piece.

Is a geo holdout viable for your business?

Carlos, our marketing science agent, can talk through your spend, channels and market footprint and give you an honest read. He will tell you if the answer is no.

Ask Carlos