Dark Wave Marketing Science Ask Carlos

Measurement

Why a media mix model needs an experiment

A model fitted to observation alone will return confident numbers whether or not the data can support them. Nothing inside it will tell you which you have.

Trevor B., Dark Wave Marketing Science ·

A media mix model always answers. Give it two years of spend and revenue and it returns a contribution for every channel, a saturation curve, and an interval around each one. It does this whether the data can support the answer or not. Producing an answer is a property of the method. It is not a verdict on the input.

Fitting is not identifying

The model observes that spend and revenue moved together and assigns credit accordingly. For that assignment to be causal, the variation in spend has to be unrelated to everything else that moves revenue.

It never is, because media budgets are not drawn at random. They are planned, and spend rises ahead of a season the business already expects to be strong. It follows a launch, answers a competitor, responds to a good quarter or a bad one. Every one of those is a reason for revenue to move that has nothing to do with the media, and each arrives in the data wearing the same shape as the media.

A model has no way to separate spend that caused revenue from spend that was scheduled in anticipation of it, and both of them look like a channel that is performing.

The priors carry more than you think

A Bayesian framework requires priors: on each channel's effect, on how long that effect carries, on where it saturates. This is a feature rather than a flaw: priors are what let a model produce something usable from a history too short to speak for itself, and they are how genuine knowledge about a business enters the estimate.

They also mean the posterior can be, in substantial part, a restatement of what was assumed. When a channel's spend barely varies, or two channels move together across the whole record, the data has little to say and the prior does the talking. The output looks identical either way: same curves, same intervals, same confident chart.

There is no field in the output labeled "how much of this came from the data."

Seasonality is the expensive one

Every model controls for seasonality, and every one of those controls is itself fitted from the same history. Where media is consistently bought into demand rather than against it, seasonal effect and media effect are entangled in the record, and the split between them is decided partly by the structure of the model rather than discovered in the data.

Get that split wrong and nothing fails loudly. Contribution simply shifts between channels, and the model goes on fitting the history as well as it did before.

The model cannot audit itself

The usual defenses measure the wrong thing.

Goodness of fit measures fit, and a model can track history closely while still allocating that history to the wrong causes, because many different allocations produce a similar curve.

Out-of-sample validation is better, and still not sufficient, because what it tests is prediction. Prediction and attribution are different claims, and a model can forecast the next quarter acceptably while being wrong about which channel is responsible. Those are precisely the errors that matter, because attribution is what the budget follows.

Nothing internal separates a model that learned the business from a model that learned the calendar. From the inside, a right answer and a wrong answer are the same object.

This is the same distinction that makes last-click a convention rather than a finding. Credit is an assignment. Cause is a counterfactual. A model that has only ever seen planned spend has never observed a counterfactual, however sophisticated the arithmetic sitting on top of it.

An experiment supplies what the history cannot

An experiment creates variation nobody planned. A holdout removes media from somewhere it would have run, or adds it where it would not have, for reasons unrelated to expected demand. The comparison that follows is a counterfactual rather than an inference drawn from correlation.

That gives you one channel, over one period, where the causal answer is known independently of the model. It is worth saying that the experiment has to be sound on its own terms first, because a test that cannot detect the effect it is looking for makes a poor yardstick for anything.

Then you ask the model to reproduce it.

That question has an answer, and it is the only external answer available. If the model's estimate for that channel lands near the experimental result, one part of it has been checked against something outside itself. If it does not, something inside is wrong, and it might be the priors, the seasonality treatment, the saturation shape, or the way effects are carried forward. Establishing which one takes real work. But you now know there is something to find, and that is a fact no diagnostic inside the model could have given you.

What this buys, and what it does not

It is worth being precise about the size of the claim, because calibration is routinely oversold.

One experiment checks one channel, over one period, at one point on its spend curve. It does not validate the model. It does not make the other channels' estimates true. It does not transfer to a channel with a different buying mechanism or a different position relative to the purchase.

What it does is convert the model from a confident instrument into a checked one, in the single place where checking was possible. Run more experiments and more of the surface becomes anchored. That is a slower and less satisfying claim than the chart implies, and it is the honest one.

A model with no experiment behind it is not necessarily wrong. It is unfalsified, which is a different condition, and the difference matters most at the moment the number is about to move a budget.

Is your model checked, or only confident?

Carlos can talk through what would have to be true for an experiment to calibrate your model, and whether your footprint supports running one at all.

Ask Carlos