A geo test can be designed correctly and still produce nothing anyone should act on. The methodology is only half the problem. The other half is procedural, and it is where most measurement quietly fails.
Degrees of freedom
The test gets designed, the data comes back, and only then does anyone settle what counts as a good result. Every choice still open at that moment is a degree of freedom, and degrees of freedom are how careful people in good faith arrive at flattering answers. Nobody has to cheat. The analysis simply drifts toward the result the room was hoping for, one defensible small decision at a time.
The counter is to memorialise the design before any outcome data exists, in a document both Finance and Marketing sign. Not a plan in someone's head and not a deck. A dated artifact that fixes the outcome definition, the treatment and control assignment and the reasoning behind it, the window and its boundaries, the effect size the design is powered to detect, and the result that would make you conclude the channel is not incremental.
Each of those is harder than it sounds, and the difficulty is where the value sits.
The outcome definition
This is usually where Finance and Marketing discover they have been running on different numbers for years. Revenue recognised when, net of what, attributed to which entity. That conversation is unpleasant, and it is better to have it in advance of a result than in the meeting where the result is presented.
The window boundaries
A modelling assumption wearing the costume of an administrative detail. Media does not stop working the day you stop paying for it, and demand does not respond the day you start. Where you cut the window encodes a belief about carryover and lag, and different defensible beliefs produce different answers from identical data. Deciding it afterwards means deciding it while looking at those answers.
Assignment and power
These are not separate decisions, though they are almost always made as though they were. Which markets you treat changes what the design can detect, so choosing markets first and checking power afterwards gets the order backwards and leaves you with a test that cannot answer the question you built it for.
The falsification criterion
The one that quietly gets negotiated away, because writing down in advance what would make you abandon a channel is genuinely uncomfortable. It is also the single line that makes everything above it credible. A test with no stated way to fail is not an experiment. It is a data collection exercise with a conclusion attached afterwards.
The document earns its keep exactly once: on the day the result is inconvenient. At that point the argument is either about what the evidence shows, or about which weeks should have been included. Only one of those conversations ends in a decision.
Why most null results are design failures
A geo test can only detect an effect larger than its noise floor. That floor is set by how many markets you have, how much they vary week to week, how long the test runs, and how large the spend change is.
If the noise floor sits above the effect you are looking for, the test cannot find the effect even when it is real. You get a wide interval spanning zero, and someone reads that as evidence the channel does not work. It is not. It is evidence the test was not built to answer the question, and the two are indistinguishable after the fact unless the power calculation was written down before the test ran.
Reading the result honestly
A geo test returns a range, not a number. If the interval is wide, the honest summary is that the effect is somewhere in that range and the decision should be made accordingly. Reporting the midpoint as though it were the answer is how a test that proved very little becomes a budget reallocation.
Concurrent activity inside the window lands in the estimate too. A promotion, a store opening, a PR moment, a competitor launch. You cannot prevent all of it. You can log it as it happens, and you can decide in advance whether a market experiencing it gets excluded, which is a different thing from deciding once you can see what excluding it would do to the result.
What the result is actually for
The most valuable use of a geo test is not the test itself. It is calibration.
A media mix model fitted only to observational data is correlational. It will produce confident numbers whether or not those numbers mean anything, because it has no way to distinguish media driving sales from media following sales. A geo experiment gives you one channel where you know the causal answer. Forcing the model to reproduce that answer, and adjusting where it does not, is what turns a fitted model into a calibrated one.
That combination is what makes continuous measurement possible. Experiments are expensive and periodic. The model carries the load in between. Neither is sufficient alone: a model without experiments drifts, and experiments without a model only tell you about the channels and periods you happened to test.
A null result is a real result. If a channel is not doing what everyone believes it is doing, that is the finding, and it is worth more than a flattering number nobody can defend.