TL;DR — “Did our marketing work?” has a right answer, but only if your data can support one. This picker asks you four yes/no questions about what you have — a comparable group, a run of history, an exact change date, a set of similar untreated units — and tells you the highest-rigor causal method you can credibly use, the one assumption it rests on, and what to do next. It walks the same identification ladder the Growth Mapping methods use: Interrupted Time Series → Difference-in-Differences → CausalImpact → Synthetic Control. No login.


The number went up. That isn’t proof.

Almost every marketing “result” you have ever seen is a before-and-after screenshot. We ran the campaign, traffic went up, the campaign worked. That reasoning quietly assumes nothing else changed in the same window — but something always does. Google shipped an algorithm update. It was the busy season. A competitor paused their ads. A press mention landed the same week.

The honest question is harder: what would have happened without the change? That invisible alternate timeline — the counterfactual — is the whole game. Everything we call “measurement” is really an attempt to estimate it. The difference between a credible marketing result and a hopeful one is entirely about how well you can reconstruct the timeline where you did nothing, and subtract it out.

Here’s the part nobody tells small businesses: you do not get to pick your measurement method by preference. You earn it with your data. A weather forecaster with one thermometer can’t model a hurricane. The rigor you can claim is capped by what you can show. This tool tells you your ceiling.


How it works

Answer four yes/no questions about your data. The picker returns the highest rung on the measurement ladder your answers support — the most rigorous causal method you can credibly stand behind — plus the single assumption that method rests on, what it needs, and a one-line next step.

It does not ask you to do any math. It tells you which method is honest given what you have. No login. The logic runs in your browser.



The identification ladder: four rungs, each earned

The methods on this picker are not interchangeable. They form a ladder. Each rung makes a stronger claim about cause, and each rung asks more of your data to back that claim. You climb as high as your evidence allows — no higher.

Rung 0 — Interrupted Time Series (ITS). You have a long run of history and you know the exact day the change went live. You model the pre-trend, then test whether the line breaks — in level or slope — at that moment. Its one load-bearing assumption: nothing else happened at the same time. That assumption is fragile, which is why ITS sits at the bottom. With no control group, anything else that landed that week wears your campaign’s clothes.

Rung 1 — Difference-in-Differences (DiD). Now you have a comparable group you did not touch — a region, a segment, a channel. You compare the change in your numbers to the change in theirs over the same window. The shared shock — seasonality, an algorithm update, the economy — hits both and cancels out. The assumption that does the work: without the change, you and your control would have moved in parallel. Stronger than ITS, because the control absorbs the confounds ITS can’t see.

Rung 2 — CausalImpact (Bayesian structural time series). With a control series and a solid history, you can do better than a single difference. Brodersen and colleagues at Google built CausalImpact precisely for this: fit a Bayesian structural time-series model on the pre-period, use the control to predict the counterfactual — what your metric would have done — and read the gap as the effect, with its uncertainty attached.1 The assumption: your control is unaffected by the change and tracks you stably. This is the workhorse of modern marketing measurement.

Rung 3 — Synthetic Control. The top rung. When no single control is a good match, Abadie and colleagues showed you can build one: take a pool of similar untreated units — other locations, clients, products — and find the weighted blend of them that best reproduces your own pre-trend.2 That weighted blend is your “synthetic twin.” After the change, the gap between you and the twin is your lift. The assumption: a weighted combination of untreated units reproduces your pre-period. It is the most credible answer a small business can give to “what would have happened anyway” — and it needs the most data to earn.

The point of the ladder is honesty. Most marketing reporting stands on a rung it hasn’t earned — claiming a clean result from a before-and-after that has no control and no counterfactual. This picker tells you which rung your data actually supports, so the confidence you put on the number matches the evidence behind it.


Knowing the method is step one. Pre-registering it is the lever.

The uncomfortable truth about the ladder is that which rung you can reach is decided before you run the campaign, not after. If you didn’t log a control group and you didn’t write down the exact change date, no method can rescue the result later. You’re stuck on “before/after,” which is no rung at all.

This is where most measurement quietly fails. The control is chosen after the fact, the date is fuzzy, and the analysis gets reverse-engineered until the number looks good. That isn’t measurement — it’s decoration.

The fix is structural. Stefan Thomke’s research on business experimentation makes the case plainly: the organizations that compound are the ones that build testing into how they operate, not the ones that test occasionally and argue about the rest.3 Applied to measurement, that means deciding your control and your counterfactual before the play runs, so the result cannot be massaged afterward.

The Hiilite platform does exactly this. Every Play pre-registers its control set at launch — the held-back segment, the donor pool, the exact go-live timestamp — so by the time you want to know if it worked, you are already standing on a high rung. The counterfactual is built into the play, not reconstructed from memory. That is the difference between a dashboard that shows what happened and a system that proves what you caused.

Read more: The Growth Mapping framework and Growth Mapping: the research behind the platform.


FAQ

How do I measure the impact of a marketing campaign?

You estimate the counterfactual — what your metric would have done if you’d done nothing — and subtract it from what actually happened. The leftover is your real effect. How you estimate the counterfactual depends entirely on your data: a held-back control group, a control series feeding a model, or a synthetic blend of comparable units. The picker above tells you which of those you can credibly use. What you should never do is compare “before” to “after” and call the difference your result, because too many other things move in the same window.

What are Difference-in-Differences, CausalImpact, and Synthetic Control, in plain terms?

Difference-in-Differences compares how much you changed to how much a similar group you didn’t touch changed over the same period — the shared seasonality and shocks cancel out, and the leftover gap is your effect. CausalImpact (built by Google researchers) uses a control series and your own history to model the timeline where the change never happened, then measures you against that prediction with honest uncertainty bands. Synthetic Control is for when no single comparison group fits: you build a “synthetic twin” out of a weighted mix of similar untreated units that matches your past, then read the gap between you and the twin after the change. They’re rungs on the same ladder — each makes a stronger causal claim and needs more data to back it.

Can’t I just compare before and after?

You can show it, but you can’t prove anything with it. A before-and-after has no way to separate your campaign from everything else that happened at the same time — the season, an algorithm update, a competitor’s move, a news cycle. That’s the lowest position on the ladder, “not provable yet.” The moment you add either a control group or a clean intervention date, you climb to a method that can actually isolate your effect. Before-and-after is a starting point, not evidence.

What data do I need to prove marketing worked?

Four things, and each one unlocks a higher rung. A run of history before the change (a baseline). The exact date the change went live (a clean break point). A comparable group you didn’t apply the change to (a control). And, for the strongest claim, several similar untreated units to build a synthetic twin (a donor pool). You don’t need all four to start — a history plus a date already gets you onto the ladder — but the more you have, the more rigorous and defensible your result. The single highest-leverage habit is logging a control group and the exact change date before every play, because that’s the part you can’t add retroactively.


About the author

William Walczak is CEO of Hiilite Creative Group (2014–present) and a PhD candidate in Interdisciplinary Graduate Studies at UBC-Okanagan, where his doctoral research — Growth Mapping: A Mixed-Method Study of Growth Hacking — examines how small businesses can apply rigorous, data-grounded growth frameworks without a data team. He holds an MBA (UBC) and an Engineering degree (Simon Fraser University), and was named Marketing Strategy CEO of the Year 2023 (BC) by CEO Monthly.

His published research includes Walczak, W., Li, E. P. H., & Nelson, S. (2024), “Logarithm: A Cinematic Exploration of Time,” Journal of Customer Behaviour.



  1. Brodersen, K. H., Gallusser, F., Koehler, J., Remy, N., & Scott, S. L. (2015). “Inferring causal impact using Bayesian structural time-series models.” The Annals of Applied Statistics, 9(1), 247–274. https://doi.org/10.1214/14-AOAS788 

  2. Abadie, A., Diamond, A., & Hainmueller, J. (2010). “Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program.” Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746 

  3. Thomke, S. (2020). “Building a Culture of Experimentation.” Harvard Business Review. https://hbr.org/2020/03/building-a-culture-of-experimentation