Sometimes the experiment was never run. The spend changed before anyone thought to design a test, or the change was forced by a budget cut, or the channel cannot be randomised at all. What remains is the record of what happened, and the question of how much can honestly be inferred from it.
This is observed incrementality. It is weaker than an experiment by construction, and it is still worth doing well, because the alternative in most companies is not a better method. It is a line chart and an assertion.
The 2x2: difference in differences
The workhorse design has four cells. A treated group and a comparison group, each measured before the change and after it. You take the change over time in the treated group, subtract the change over time in the comparison group, and call the remainder the effect.
The subtraction is doing something specific and worth being precise about. A simple pre versus post comparison attributes everything that changed to the intervention, including seasonality and the general direction of the business. A simple treated versus untreated comparison attributes every pre-existing difference between the two groups to the intervention. Differencing twice removes both: anything that affected both groups equally over the period drops out, and anything that made the groups persistently different drops out as well.
What does not drop out is anything that affected one group differently over the period. That residual is the entire vulnerability of the method.
Parallel trends is the whole assumption
Difference in differences is valid when the two groups would have moved in parallel had nothing happened. Not identical levels, parallel movement. Every claim the design makes rests on that, and it cannot be tested in the period where it matters, because that is precisely the period where one group was treated.
What you can do is examine whether the groups moved in parallel for a long stretch beforehand. That is evidence, and it is the right thing to show a sceptical reader, but it is not proof. Groups can track each other for a year and diverge for reasons unrelated to the intervention.
The failure that matters most is when the treatment was not assigned arbitrarily. If spend was increased in the markets that were already accelerating, or cut in the ones that were already struggling, the groups were on different paths before anything happened and the estimate absorbs that difference. The intervention was chosen in response to the outcome, which is the one condition under which this design cannot recover.
Before trusting a difference in differences result, ask how the treated group was selected. If the answer involves the outcome you are now measuring, the number is not what it appears to be.
Time-based holdouts: the 2x2 with a cell missing
When there is no comparison group at all, what remains is turning spend up and down over time and watching the outcome move. Alternating periods, sometimes switched back and forth repeatedly.
Structurally this is the same design with the control column removed. The comparison is between now and earlier, so the counterfactual is the past, and every other thing that changed between now and earlier is inside the estimate. Seasonality, competitor activity, promotions, pricing, macro conditions, and your own other channels are all in there, and the design has no way to separate any of them from the media effect.
Carryover makes it harder. Media bought in an on period keeps working into the following off period, which blurs the boundary between conditions and biases the measured difference downward. The obvious fix is longer periods, but longer periods mean fewer switches, and fewer switches mean less power. The two pressures work against each other, and where they land decides whether the test can detect anything at all.
Switching repeatedly rather than once helps, because it makes it less likely that a single coincident event explains the whole pattern. It does not create a control group. It makes the absence of one less damaging.
When observed methods are the right call
They earn their place when the change has already happened and the alternative is no estimate at all, when there is a long clean pre-period to establish that the groups moved together, and when the treatment was not selected on the basis of the outcome.
They are the wrong call when someone wants certainty. An observed estimate comes with an identifying assumption attached, and the honest presentation states that assumption out loud rather than burying it in an appendix. A result presented without its assumption is not a more confident result. It is a less honest one.
The strongest use of an observed estimate is as a prior to be tested, not a conclusion to be acted on. It tells you where an experiment would be worth running. Where the conditions allow it, run the experiment.