Research · Measurement Method   Last updated June 2026

What would have happened anyway?

Incrementality is the share of conversions your marketing actually caused — not the ones it merely accompanied. We measure the gap between what happened and what would have happened if you had done nothing. That gap is the only number that proves your marketing worked.

A report tells you the number went up. We tell you whether you’re the reason it did.

The Ghost Line lift chart A line chart comparing two paths over time. A dotted slate line shows what would have happened anyway as a flat baseline forecast. A solid orange line rises well above it, representing actual results. The shaded area between the two lines is labelled “the lift”. A small charcoal pill annotation reads “What would have happened anyway” and points at the dotted line. the lift What would have happened anyway start now
The Ghost Line: actual results (solid orange) above the do‑nothing forecast (dotted). The shaded gap is the lift. +38% incremental lift

Estimate: sizing the prize before you spend a dollar

We’ll tell you the floor, the likely, and the ceiling — and we’ll tell you which one we’d bet on.

Forecast estimate range band A horizontal range from a conservative floor on the left to an optimistic ceiling on the right, with the most likely value marked by an orange tick near the middle. Likely Floor Conservative Ceiling Optimistic
An honest forecast is a band — floor, likely, ceiling — not a single number.

We don’t promise 10×. Honest forecasting comes as a conservative / expected / optimistic band, not a single number. It’s grounded in your reality: search demand, your close rate, what a customer is worth to you, and how much headroom you have left — you can’t capture demand that isn’t there.

Returns aren’t a straight line. First effort pays back fastest, then it flattens — we plan for the curve, not infinite growth. The estimate is what lets us rank plays before spending: high payoff, low cost, do‑able with your team.

A worked estimate band: one play, three scenarios
ScenarioAssumptionModelled outcome
Conservative (floor)Low demand capture, average close rate+12% qualified enquiries
Expected (likely)Realistic capture given current headroom+28% qualified enquiries
Optimistic (ceiling)High capture, strong execution, low competition+45% qualified enquiries

Illustrative figures to show the shape of an honest range — your numbers are modelled from your own economics in the tool below.

Open the SEO ROI calculator →

Measure: proving it moved

A counterfactual is the estimate of what would have happened without the marketing — the do‑nothing baseline. Because you can’t observe both outcomes for the same customer, you approximate it with a randomized control, a geo holdout, or a synthetic‑control model.

“It went up” is not proof. Seasonality, a Google update, a competitor’s move, or a concurrent sale can all push the line up without your play doing anything. So we build a counterfactual, then measure the gap — and we lock the baseline before the play launches, not after.

Two honest moves do the work: hold something comparable as a control (an untouched region, page, or channel), and check the forecast against the pre‑period so it’s trustworthy.

The Ghost Line: counterfactual baseline vs measured lift A line chart. A dotted slate forecast line continues the pre-period trend as the counterfactual baseline. A solid orange line shows the actual outcome, rising above the forecast after launch. A vertical charcoal dashed marker labelled “baseline locked here” sits at the launch point, dividing the pre-period from the measured-lift region. The shaded area between the actual line and the forecast, to the right of the marker, is labelled “measured lift”. Outcome metric Time PRE-PERIOD MEASURED LIFT baseline locked here (launch) measured lift forecast (counterfactual) Actual Counterfactual baseline Measured lift
We lock the baseline at launch, then measure the gap that opens after.
We lock the baseline before the play launches, so the result can’t be reverse‑engineered to flatter us.

Open the incrementality calculator →

Further reading: Lewis & Rao (2015), The Unfavorable Economics of Measuring the Returns to Advertising, Quarterly Journal of Economics.

The Hiilite Measurement Ladder

We don’t claim more certainty than your data can pay for — we tell you which rung of the ladder the proof is standing on.

There’s no single right way to measure — there’s a ladder, and we climb to the highest rung your data can support. Higher rung, stronger proof.

The Hiilite Measurement Ladder A four-rung vertical ladder showing causal measurement methods from bottom to top, each with a stronger strength-of-proof rating. Bottom rung: Before vs After (ITS), one star. Next: Treated vs Untouched (DiD), two stars. Next: Modelled Forecast (CausalImpact), three stars. Top rung: Synthetic Control, four stars. A rising arrow alongside indicates increasing strength of proof. STRENGTH OF PROOF Synthetic Control Weighted twin of untreated units ★★★★ Modelled Forecast CausalImpact — predicted vs actual ★★★ Treated vs Untouched Difference-in-Differences (DiD) ★★ Before vs After Interrupted Time Series (ITS)
The four rungs at once: Before vs After (ITS, ★), Treated vs Untouched (DiD, ★★), Modelled Forecast (CausalImpact, ★★★), Synthetic Control (★★★★).
Rung 1 — Before vs after (Interrupted Time Series) · ★

The simplest read: did the line change at the moment you acted? Trustworthy only if nothing else big changed at the same time. Use this when you have clean before/after data and no competing events.

Rung 2 — Treated vs untouched (Difference-in-Differences) · ★★

Run the play in one place, hold a comparable place steady, and compare the change in each. Use this when you have a matched region, page, or segment you can leave alone.

Rung 3 — Modelled forecast (CausalImpact, our default) · ★★★

Build a statistical “what would have happened” line from related signals, then measure the gap. Pioneered by Google’s research team. Use this when you have history and correlated control series but no clean holdout.

Rung 4 — Synthetic control · ★★★★

When there’s no perfect twin, we build one from a blend of comparison cases — the gold standard for a single business. Use this when the result has to stand up to real scrutiny.

Find out which rung your data supports →

Why the same play lands differently

There’s no average business, so we don’t sell you an average result — we estimate yours.

A play that doubled leads for one client might do little for another. That’s not failure, it’s context.

Heterogeneity diptych: the same play, two very different lifts Two minimalist line-icon vignettes on white. Left: a small storefront with an orange play badge above and a Ghost Line chart showing a wide orange lift gap. Right: a small office with the same orange play badge above and a Ghost Line chart showing a narrow lift gap. A caption strip lists the context factors that drive the difference: headroom, competitiveness, margin, loyalty, execution. wide lift Local shop · lots of headroom narrow lift Saturated B2B · little headroom WHAT MOVES THE GAP headroom competitiveness margin loyalty execution
Same play, two businesses, two very different gaps. The factors that move it: headroom, competitiveness, margin, loyalty, execution.
What’s different about you?

Tick the factors that favour you — the illustrative expected‑result band below widens with more advantage.

expected result

More advantages → a tighter, higher expected band. Fewer → a wider, more uncertain one.

What changes the outcome: headroom left to grow, how competitive your search results are, your pricing and margins, customer loyalty, and whether your team can execute the play to a high standard. We don’t report a single industry “average lift” and pretend it’s your number — we estimate your expected result given your situation. That’s why we’re skeptical of “this one trick works for everyone” marketing: recommendations come with the conditions under which they hold.

See the full Growth Mapping framework →

Borrowing strength: new clients start ahead

You don’t start from zero — you start from everything we’ve learned on businesses like yours, then your own numbers take over.

A brand‑new client has almost no history, so a naïve approach guesses wildly. We start your estimate from everything we’ve learned across many similar businesses, then let your real results pull it toward your truth.

Every play we measure makes the next client’s forecast smarter — the system compounds. It’s the opposite of a blank slate every engagement.

Borrowing-strength convergence toward your value A field of faint grey dots labelled “other clients” surrounds a pooled-average marker. One orange dot, mid-path, slides rightward along a thin arrow toward a “your value” target on the right. other clients pooled average your value
Your forecast starts near the pooled average of similar businesses, then your own results pull it to your value.

Glossary

The measurement vocabulary, one dictionary‑clean definition each.

Incrementality
The share of conversions your marketing actually caused, not the ones it merely accompanied.
Counterfactual
The estimate of what would have happened without the marketing — the do‑nothing baseline.
Causal attribution
Crediting an outcome to a cause that actually produced it, proven against a baseline rather than assumed.
Holdout
A randomly withheld group that sees no marketing, used as the control to measure lift.
Geo‑test
An experiment that runs a campaign in some regions and not others, then compares the outcomes.
MMM (marketing mix modeling)
A top‑down statistical model estimating each channel’s contribution from historical spend and sales.
MTA (multi‑touch attribution)
A method that distributes conversion credit across the touchpoints a converter saw.
Synthetic control
A counterfactual built from a weighted blend of comparison cases when no single twin exists.
Difference‑in‑differences (DiD)
Comparing the before/after change in a treated group against the change in an untouched one.
Lift
The measured gap between the actual outcome and the counterfactual baseline.
Statistical significance
Confidence that a measured difference is real and not just noise.
Attribution window
The time period after exposure within which a conversion is credited to a marketing touch.

MMM vs incrementality vs MTA vs A/B

Use MMM to plan the mix; use incrementality to validate it.

The canonical reference block — four methods across identical columns.

Four ways to measure marketing, compared
Method Unit of randomization What it measures Best for Main limitation Privacy‑safe?
Incrementality testing Users, regions, or time (randomized control / holdout) Causal lift — conversions your marketing actually caused Validating that a specific campaign truly worked Needs a held‑back group; forgoes some reach Yes — geo/time designs need no user tracking
MMM (marketing mix modeling) None — observational, modelled from history Each channel’s estimated contribution to sales Budget‑level, top‑down strategy and mix planning Correlational; needs long, clean history Yes — uses aggregate spend and sales
Multi‑touch attribution (MTA) None — observational, path‑based Credit shared across touchpoints on the path Operational channel reporting at the user level Assumes everyone who clicked was influenced; over‑credits paid channels No — relies on user‑level tracking
A/B test Users (randomized at the variant level) Causal effect of one change vs another On‑site changes: pages, creative, flows Limited to what you can randomize on‑site Yes — first‑party, no cross‑site tracking

FAQ

How do I know if my marketing is actually working?

You measure incrementality — the additional sales your marketing caused that would not have happened anyway. The honest test is a controlled comparison: hold back a randomized group (a holdout or geo‑test) and compare exposed vs unexposed outcomes. The gap is your causal lift; everything else is correlation a dashboard can’t separate from luck or seasonality.

What is incrementality in marketing?

Incrementality is the share of conversions a marketing activity caused rather than merely accompanied. It answers “what would have happened anyway?” by comparing a treated group against an unexposed control. Attribution credits touchpoints on the path to purchase; incrementality isolates the true causal effect against a counterfactual baseline.

What’s the difference between incrementality and attribution?

Attribution distributes credit across the touchpoints a converter saw and assumes everyone who clicked was influenced. Incrementality asks whether they would have converted without the ad at all. Attribution measures correlation along a path; incrementality measures causation against a control group — and frequently shows attribution overstates paid‑channel impact (a finding documented by Lewis & Rao, 2015).

Is MMM or incrementality testing better for a small business?

Use both for different jobs. Marketing mix modeling (MMM) is a top‑down, privacy‑safe model that estimates each channel’s contribution from historical spend and sales — good for planning the budget mix. Incrementality testing is a bottom‑up controlled experiment that proves causal lift for a specific campaign. Use MMM to plan the mix; use incrementality to validate it.

What is a counterfactual, and why does it matter?

A counterfactual is the estimate of what would have happened without the marketing — the do‑nothing baseline. Because you can’t observe both outcomes for the same customer, you approximate it with a randomized control group, a geo holdout, or a synthetic‑control model. Causal measurement is only ever as good as its counterfactual.

Can I measure incrementality without a data team or tracking pixels?

Yes. Three lightweight designs work without analysts: a geo holdout (pause ads in matched regions), a time‑based on/off test (scheduled blackouts), and a customer holdout (withhold a campaign from a random 10%). Each yields a measurable, financials‑bound lift number with no attribution model or tracking pixel required.

How long does an incrementality test take?

Most lightweight geo or holdout tests run for one to two purchase cycles, and cost mainly the foregone reach in the held‑back group. The slower your conversions, the longer the window needed to reach statistical confidence — so the right duration depends on your sales velocity and how large a difference you need to detect.

Does attribution overstate how well my ads work?

Often, yes. Attribution credits any touchpoint a converter saw, including people who would have bought anyway, so it tends to over‑credit paid channels — especially retargeting and branded search. Incrementality testing strips out that baseline and reports only the lift your marketing actually caused, which is frequently lower than the attributed number.

From doctoral research, in the product every day

We measure marketing the way a scientist would — including being willing to tell you it didn’t work. Every play we run declares up front what it should move, how we’ll prove it, and what would count as failure. This rigor comes from doctoral research and runs in the product every day.

William Walczak presenting at a social-media marketing convention
Presenting on measurement at a marketing convention.
William Walczak leading a client strategy meeting
Leading a client growth‑mapping session.
William Walczak, MBA, PhD candidate, CEO of Hiilite Creative Group Inc.
William Walczak — MBA, PhD candidate (UBC‑Okanagan), CEO, Hiilite Creative Group Inc.

References & further reading

  • Lewis, R. A., & Rao, J. M. (2015). The Unfavorable Economics of Measuring the Returns to Advertising. The Quarterly Journal of Economics, 130(4), 1941–1973.
  • Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.
  • Teece, D. J. (2007). Explicating dynamic capabilities: the nature and microfoundations of (sustainable) enterprise performance. Strategic Management Journal, 28(13), 1319–1350.
  • Google — Meridian: open‑source marketing mix modeling. github.com/google/meridian
  • Meta — GeoLift: open‑source geo‑experiment measurement. github.com/facebookincubator/GeoLift

Ready to find out whether your marketing is the reason?

One conversation to map what to measure, and how we’d prove it.