The Measurement Ladder: match your proof to your data
Most measurement advice is either useless (“just look at last-click”) or impossible (“run a randomised experiment”). The Measurement Ladder is the in-between: a staged marketing measurement framework that tells you what kind of proof you can honestly claim given the data you actually have — and what to add next.
By William Walczak — CEO, Hiilite Creative Group Inc.Why one measurement method never fits everyone
A solo founder running $800/month in ads and a 40-person company spending six figures a quarter do not have the same measurement problem. The founder cannot run a clean holdout test — there isn’t enough volume for the maths to settle. The larger company shouldn’t be staring at last-click attribution when it has the budget to run a real counterfactual. Yet most advice treats “how should I measure marketing?” as one question with one answer.
The honest answer depends on two things: how much data you generate, and how much rigour the decision in front of you deserves. The Measurement Ladder pairs those two together. Lower rungs give you fast, cheap, directional reads. Higher rungs cost more — in traffic, time, and discipline — but tell you something closer to the truth: did this marketing actually cause the result, or would it have happened anyway? That “would it have happened anyway?” question is incrementality, and it’s what every rung is quietly climbing toward.
The four rungs, from directional to causal
Think of the ladder as four rungs. You don’t skip steps for prestige, and you don’t stay on the bottom rung once you’ve outgrown it. Each rung answers a slightly harder question with slightly better proof.
- Rung 1 — Attribution + directional reads. Last-click, UTM tracking, “how did you hear about us?” on the booking form. It’s correlational and it over-credits whatever the customer touched last, but at low volume it’s often all you can run — and a rough signal beats no signal. Use it to spot obvious winners and obvious duds, not to settle close calls.
- Rung 2 — Multi-touch + before/after comparisons. Looking across the whole journey and comparing periods (the month you launched a campaign vs. the month before). Better context than last-click, but still confounded by seasonality, promotions, and everything else moving at once.
- Rung 3 — Marketing Mix Modelling (MMM). A statistical model that estimates each channel’s contribution to sales while controlling for season, price, and baseline demand. Privacy-proof and channel-wide, but it’s a model of the past, not a controlled test — correlation cleaned up, not proof.
- Rung 4 — Randomised holdouts & geo-tests. Deliberately withhold marketing from a random group (or a set of regions) and compare. This is the only rung that produces a true counterfactual — a real measure of incremental lift. It needs volume and patience, which is exactly why it sits at the top.
The rule: don’t demand proof you can’t run — or settle for proof you’ve outgrown
The Measurement Ladder has two failure modes, and they pull in opposite directions. The first is over-reaching: demanding a randomised experiment to justify a $300 test, or refusing to act until you have “clean” causal data you’ll never be able to collect at your size. That’s paralysis dressed up as rigour. If you can’t run a holdout, the right move is a confident Rung 1 or 2 read, clearly labelled as directional.
The second failure mode is under-reaching: a business with plenty of volume and budget still pointing at last-click to make big spend decisions. If you can run a holdout, settling for attribution is leaving real money on the table — you’re crediting channels that may be taking credit for sales you’d have won anyway. The discipline is simple: climb to the highest rung your data can support for the size of the decision, and no higher. Our full method on how these methods relate lives in how we measure growth.
Where are you on the ladder? A 60-second self-check
Answer these honestly. They’re about volume and decision stakes, not ambition.
- Fewer than ~30–50 conversions a month? You’re on Rung 1–2. Holdouts won’t reach significance — use attribution to spot the obvious, and make calls directionally. Don’t apologise for it.
- Hundreds of conversions and several channels running at once? You’re ready for Rung 3 (MMM) to untangle what’s really driving sales beyond the last click.
- Thousands of conversions, or multiple regions/markets? Rung 4 is open to you. For your biggest spend decisions, run a holdout or geo-test and measure actual lift.
- Making a high-stakes call (cutting a channel, doubling a budget)? Push one rung higher than usual for that specific decision — the proof should match the stakes, not just the data.
- Making a low-stakes call (a $200 creative test)? Stay low. A directional read is the right tool; a randomised experiment is overkill.
Put the ladder to work with the free tools
You don’t need a data team to start climbing. If you’re sizing whether a holdout would even register, the incrementality calculator shows the lift you’d need to see and whether your volume can detect it — the fastest way to know if Rung 4 is realistic yet. If you’re weighing channels and not sure which method fits, the measurement method tool recommends a rung for your situation, and the SEO ROI calculator helps you reason about return before you commit spend.
The ladder isn’t academic for its own sake — it’s about not lying to yourself. A clearly-labelled directional read you can act on this week beats a perfect experiment you’ll never run. Pick your rung, claim only the proof it gives you, and climb when your data earns it. If you want the vocabulary behind all of this, the measurement glossary defines every term, and the MMM vs. incrementality vs. attribution comparison walks the three big methods side by side.
Keep reading
Questions, answered
What is the Measurement Ladder?
It’s a staged marketing measurement framework with four rungs — attribution, multi-touch/before-after, marketing mix modelling, and randomised holdouts. Each rung gives stronger proof but demands more data. The idea is to use the highest rung your conversion volume can support, sized to how big the decision is.
Do I need to run experiments to measure my marketing well?
No. Randomised holdouts are the top rung and produce true causal proof, but they need real volume to work. If you have fewer than roughly 30–50 conversions a month, a clearly-labelled directional read from attribution is the honest, correct choice. Demanding a perfect experiment you can’t run is just paralysis.
When should I move up to MMM or a holdout test?
Move to MMM (Rung 3) once you have hundreds of conversions across several channels and last-click can no longer untangle them. Move to holdouts or geo-tests (Rung 4) when you have thousands of conversions or multiple markets — and especially for high-stakes decisions like cutting a channel or doubling a budget.
More from this series
Seven plain-English pieces on measuring marketing and turning tests into durable capability:
Want to know whether your marketing is the reason?
One conversation to map what to measure and how we’d prove it.