Dark Wave Marketing Science Ask Carlos

Measurement

Alternatives to geo testing

Not every advertiser has the geographic variation a holdout needs. Three designs remain, and they are not interchangeable.

Dark Wave Marketing Science

A geo holdout needs geography to work with. Plenty of advertisers do not have it: one metro, or national buying with no market-level control, or an outcome that cannot be resolved to a location. The question then becomes which of the remaining designs is worth running, and what each one can honestly claim.

The useful way to sort them is by who holds the randomisation. In the first, you do. In the second, the platform does. In the third, nobody does. That single ordering predicts most of what follows, including how much of the result you are in a position to verify yourself.

Ghost ads and PSA holdouts

Here the randomisation happens in the ad server. A slice of the addressable audience is held out. In a PSA design they are served an unrelated public service ad instead of yours. In a ghost ad design they are served nothing, but the server logs the impression your ad would have won, so both groups are defined by the same eligibility rather than by whether an ad happened to appear.

That last point is the reason this design is strong. The usual failure of user-level measurement is that the exposed group and the unexposed group differ in ways that caused the exposure. People who see your ad are people the targeting selected, and targeting selects for propensity to buy. Ghost ads compare people the system wanted to reach against other people the system wanted to reach, which is a much narrower gap to argue about.

The constraints are practical rather than conceptual. It requires an ad server or DSP that supports the design, which rules out most walled gardens where the inventory cannot be instrumented this way. The counterfactual impression is itself an inferred quantity, so the quality of the control depends on the quality of that inference. And it measures the channel it runs in. Effects that show up in another channel, in retail, or in brand response over a longer horizon are outside what it can see.

Platform lift studies

Every large platform offers one. The randomisation happens inside the platform, on its users, against its outcome data, and the result comes back as a finished number.

They are fast, cheap, and often free, and for a first read on a single channel they can be genuinely useful. The structural problem is not the statistics. It is that the party being measured designs the test, defines what counts as exposure, chooses the outcome window, computes the result, and delivers it. Every one of those is a defensible choice with a direction, and you are not in a position to audit any of them.

Two consequences follow that catch people out. Lift measured by different platforms is not comparable, because each is using its own definition of exposure and its own attribution window against its own view of conversions. And lift figures cannot be added together. Audiences overlap, so two platforms can each honestly claim the same incremental conversion, and a portfolio built by summing them will describe a business that does not exist.

A platform lift study answers the question "did this platform's advertising work, as this platform defines working." That is a narrower claim than most decks make with it.

Matched panel

The third design does not randomise anything. You observe who was exposed and who was not, then match individuals in the exposed group to individuals in the unexposed group on their pre-period behaviour, and compare what happens next.

It is the most widely available of the three, because it needs data rather than infrastructure. Anywhere you have user-level records, a logged-in base, a CRM, or a panel provider, you can run it. For channels that cannot be instrumented any other way, it is often the only option on the table.

The weakness is the one that matching cannot fix. In a ghost ad design, both groups were selected by the same system for the same reason and one arm simply did not get served. In a matched panel, exposure happened for reasons you did not control and mostly cannot observe. You match on the behaviour you can see, and the thing that drove exposure may be the thing you cannot see. If the targeting found people who were already close to buying, matching on last quarter's purchases does not remove that, it just makes the two groups look similar in the dimension you happened to measure.

There is a second problem specific to media. Your unexposed group is unexposed in your data. Whether they actually saw the advertising somewhere you cannot observe is a different question, and every one of them who did pushes the measured effect toward zero.

This puts matched panel in a different category from the two above it. Those are experiments with a compromised or borrowed counterfactual. This is an observational estimate wearing experimental clothing, and it should be read with the caution that implies. The methods for doing that honestly are a subject of their own.

What none of them replace

All three answer channel-level or platform-level questions. None of them produces a market-level read across your whole mix, which is what a budget conversation actually requires, and none of them covers the parts of the business that do not run through an ad server: organic demand, retail, brand, offline.

That gap is why the model matters. Experiments of any kind, geo or otherwise, are periodic and narrow. Their most valuable use is calibrating a model that runs continuously and covers everything. A design that gives a trustworthy causal read on one channel is worth more as a calibration point than as a headline.

The right question is not which test is best. It is which question you need answered, and what the cost of being wrong about it is. Where geography is available, geo design remains the strongest option for a market-level answer.

Which design fits your setup?

Carlos can talk through your channels, footprint and data, and say plainly which of these is viable for you and which is not.

Ask Carlos