Marketing measurement

Marketing Effectiveness: How to Measure What Actually Works

The methods, the metrics that matter, the maths worked through, and the budget decisions each one is meant to drive.

AL
Aryma Labs
Aryma Labs
19 min read

Definition

Marketing effectiveness is how well marketing activity delivers against business objectives, measured across both short-term response and long-term brand building. It is not a single metric but a measurement system, and the most common failure is optimising the part that is easy to measure while the part that compounds goes unmeasured.

Marketing effectiveness is the degree to which marketing activity produces the business outcomes it was funded to produce: incremental revenue, profit, customer acquisition and brand equity. It is not a single metric. It is a measurement system that separates the sales marketing caused from the sales that would have happened anyway, and holds that judgement steady across cycles.

Most guides on this topic hand you a list of KPIs and stop. The harder problems sit underneath the list: which method is appropriate for your spend, how wide the uncertainty around your estimate really is, what to do when three methods disagree, and how much of your return lands outside the window you are reporting on. This piece works through those.

What marketing effectiveness actually measures

Three ideas get conflated in most effectiveness conversations, and pulling them apart is most of the job.

  • Output is what marketing produced. Impressions, clicks, sessions, leads, reach.
  • Outcome is what changed in the business because of it. Incremental revenue, new customers, gross profit, penetration, brand equity.
  • Efficiency is what each unit of output or outcome cost. CPM, CPC, cost per lead, cost per acquisition.

Effectiveness lives in the second category. A campaign can be extremely efficient and completely ineffective. Cheap clicks delivered to people who were already going to buy are, in causal terms, worth nothing at all. The counterfactual, meaning what would have happened without the spend, is the only honest benchmark.

Marketing effectiveness vs marketing efficiency

The distinction is not academic. It determines which number wins an argument in a budget meeting.

Marketing effectiveness Marketing efficiency
Question answered Are we doing the right things? Are we doing things at the right cost?
Typical measures Incremental revenue, incremental profit, penetration, share, brand equity CPA, CPL, CPM, cost per session
Time horizon Quarters and years Days and weeks
Method Causal: modelling and experiments Arithmetic: cost divided by volume
Failure mode Executing the wrong strategy very well Optimising a channel that never contributed
Primary audience CFO, board, exec team Channel and performance teams

Efficiency metrics are easy to compute, which is exactly why they dominate reporting. They are also easy to game. Any channel can improve its cost per acquisition by narrowing its targeting onto people already in the market, which lowers cost and lowers incremental contribution at the same time.

The pressure to answer the effectiveness question is not going away. Gartner's 2025 CMO Spend Survey reports that marketing budgets have flatlined at 7.7% of overall company revenue in 2025, unchanged from 2024. When the budget stops growing, reallocation becomes the only remaining source of growth, and reallocation requires a defensible view of what each dollar is actually doing.

Why marketing effectiveness is hard to measure

There is a wide gap between how confident marketers feel and what they can actually evidence. Nielsen's research accompanying The Marketing ROI Blueprint found that 85% of marketers express confidence in their ability to measure ROI, while only 32% actually measure ROI holistically across traditional and digital media channels. Confidence is not the constraint. Method is.

Five structural problems sit behind that gap.

  1. Fragmented data. Spend lives in platform interfaces, revenue lives in the finance system, and the two are rarely joined at a common grain of date, geography and product.
  2. Signal loss. Consent requirements, browser and device restrictions and platform-level aggregation mean user-level journeys are increasingly incomplete, and incompletely at random.
  3. Every platform marks its own homework. Ad platforms report conversions they observed, using their own attribution windows, with no view of what would have happened had the ad not served. Summing platform-reported revenue routinely exceeds actual revenue.
  4. Lag and carryover. Advertising keeps working after the flight ends. A window that closes too early systematically understates effect.
  5. Correlation dressed as causation. Spend usually rises alongside seasonality, promotion and distribution. Without controls, a model will happily attribute all of it to media.

The marketing effectiveness metrics that matter

A useful way to organise marketing effectiveness metrics is a four-level ladder, where each level is only meaningful if the level below it is read in context.

  1. Activity metrics. What the team did: campaigns shipped, emails sent, content published.
  2. Operational metrics. How the media performed: impressions, reach, frequency, click-through rate, view rate.
  3. Outcome metrics. What the customer did: conversions, leads, orders, repeat purchase, retention.
  4. Business metrics. What changed on the P&L: incremental revenue, incremental gross profit, customer acquisition cost, customer lifetime value, market share.

Effectiveness is settled at level four. Levels one to three are diagnostics that explain a level-four result, and they become vanity metrics the moment they are reported as the result itself.

Core formulas

Metric Formula What it answers Where it misleads
ROAS Attributed revenue / media spend Revenue returned per unit of spend Uses attributed, not incremental, revenue and ignores margin
iROAS Incremental revenue / media spend Revenue that would not have occurred otherwise Requires a valid control group or model
ROMI (Incremental gross profit - media spend) / media spend Profit generated per unit of spend Highly sensitive to the margin assumption used
CAC Total acquisition cost / new customers acquired Cost of buying one customer Blended CAC hides channel-level variation
CLV Average order value x purchase frequency x lifespan x gross margin Value of a customer over the relationship Rests entirely on the retention assumption
CLV:CAC CLV / CAC Whether acquisition economics work Meaningless without a stated payback period

ROI and ROMI are not the same calculation. ROI as commonly reported divides revenue by spend. ROMI works in gross profit and subtracts the media cost, so it answers the question a CFO is actually asking: did this investment return more than it consumed?

The metrics to stop reporting

  • Impressions and reach as standalone success measures.
  • Social followers and engagement rate, absent any link to purchase behaviour.
  • Email open rate, now materially distorted by privacy proxies.
  • Total attributed conversions summed across platforms.
  • Last-click revenue presented as channel contribution.

None of these are useless as diagnostics. All of them are misleading as answers.

A worked example: measuring incremental effect end to end

Every guide defines the ROMI formula. Almost none run one. Here is a complete geo holdout read, with illustrative but internally consistent figures, so you can see where the judgement calls actually sit.

Design. Eighty matched markets, split into 40 test markets where paid social continues and 40 control markets where it is switched off entirely. The test runs for eight weeks. Markets were matched on twelve months of pre-period revenue, and the pre-period test-to-control revenue ratio was 1.00.

Line Value
Control market revenue, 8 weeks $2,000,000
Test market revenue, 8 weeks $2,300,000
Pre-period test:control ratio 1.00
Counterfactual revenue for test markets $2,000,000
Incremental revenue $300,000
Media spend in test markets $150,000
Incremental ROAS 2.0
Gross margin 55%
Incremental gross profit $165,000
ROMI (165,000 - 150,000) / 150,000 = 10%

Three things are worth noticing.

The platform number was almost twice the truth. Over the same eight weeks the ad platform reported $520,000 of attributed revenue in those markets, implying a ROAS of 3.47. The measured incremental ROAS was 2.0. The gap is not fraud. It is the platform correctly counting conversions it observed, with no way of knowing which of them would have happened regardless.

A healthy ROAS produced a thin ROMI. An incremental ROAS of 2.0 sounds strong until margin and the cost of the media itself are applied, at which point the activity returns ten cents on the dollar. Whether that clears the hurdle depends on your cost of capital and on repeat purchase behaviour this window cannot see.

The point estimate is not the answer. Market-level revenue is noisy. Suppose the 95% confidence interval on incremental revenue runs from $160,000 to $440,000. Incremental ROAS is then somewhere between 1.07 and 2.93, and ROMI between roughly -41% and +61%. That range is the honest result. Reporting "ROMI was 10%" without it invites a decision the evidence cannot support.

This is why minimum detectable effect matters before a test runs rather than after. If your market-level variance means the smallest lift you can reliably detect is 12%, and the effect you expect is 5%, the test will return a null result regardless of whether the channel works. Fix that with more markets, a longer run, or a larger spend differential. If none of those are available, do not run the test.

Marketing effectiveness measurement methods compared

Four families of method dominate. They answer genuinely different questions, and the common advice to "run all three" ignores that most organisations cannot support all three.

Method What it can prove Data required Cadence Where it fails
Marketing mix modelling (MMM) Contribution, saturation and carryover for every channel, including offline 2 to 3 years of weekly spend, sales and control variables Quarterly Thin or collinear spend variation, too few observations, short histories
Multi-touch attribution (MTA) How observed digital journeys distribute credit User-level, cross-device event data with consent Daily Consent gaps, walled gardens, offline channels, and it measures correlation not causation
Incrementality and geo testing Causal lift for one channel or tactic over one window Geographic or randomised split with clean revenue reporting Per test Costly in forgone spend, narrow in scope, needs statistical power
Calibrated MMM Channel contribution anchored to experimental ground truth Both of the above, plus the discipline to keep testing Continuous Organisational rather than technical: it needs a standing test programme

The fourth row is where credible practice has landed. Google's open-source Bayesian framework Meridian documents this explicitly: calibrating MMM with incrementality experiments across channels, modelling reach and frequency so video planning links to outcomes, and controlling for organic demand using search query volume to improve lower-funnel accuracy. Experiments and models are not rival methods. Experiments supply the priors that keep the model honest.

It is also worth saying plainly where MMM is the wrong tool. If you spend under roughly $1m a year, run three channels, and hold eighteen months of history, an MMM will produce coefficients with intervals so wide that any allocation decision drawn from them is guesswork with a decimal point attached. Experiments will serve you far better.

Which measurement method should you actually use

Method selection should follow spend, channel count and data history, not vendor preference. The rule of thumb below is practitioner judgement rather than a published standard, but it reflects where each method starts earning its cost.

Annual working media spend and years of clean weekly data Under $1m or under 2 years of data $1m to $10m 2 to 3 years of weekly data Above $10m 3+ years, many channels Experiments first Run geo holdouts and matched-market tests. Spend rarely varies enough for MMM to identify channel effects. Platform numbers stay directional only. Experiments, then MMM Make incrementality testing the backbone. Add a Bayesian MMM once you hold two to three years of clean weekly data, and calibrate it against the tests you have run. MMM as system of record Model continuously and calibrate against a standing experiment programme. Use platform reporting for in-flight pacing, never for board-level contribution. The constant across all three Platform-reported ROAS is a pacing signal for in-flight optimisation, not a measure of incremental contribution.
Choosing a marketing effectiveness method by spend and data history. The thresholds are practitioner judgement, not a published standard.

B2B and long sales cycles are a different problem

The guidance above assumes conversion follows exposure within weeks. In B2B, where a first touch and a closed deal can sit six to eighteen months apart, that assumption breaks completely. Three adjustments help.

  • Model pipeline creation, not closed revenue, as the near-term outcome variable, and validate the pipeline-to-revenue conversion rate separately.
  • Match the measurement window to the observed cycle length, taken from your own closed-won records rather than an assumed quarter.
  • Use account-level and market-level holdouts rather than user-level splits, because the buying unit is a committee, not a cookie.

Related product

MMM Synapse

Turn years of legacy MMM decks into a conversational, encrypted knowledge base you can query in plain language.

See MMM Synapse

The long half of marketing effectiveness

The largest hole in most effectiveness reporting is temporal. A majority of the return on media does not arrive inside the window most teams report on.

Analysis published by Think with Google, drawing on the Google and WARC research, found that returns on media investments in the first four months are equal to the returns generated across the subsequent 20 months. In profit terms, short-term profit ROI averages £1.87 for each £1 of investment, rising to £4.11 once sustained long-term effects are included.

Most of the return arrives after the reporting window closes Cumulative profit return per £1 of media investment, per Google and WARC analysis £1 £2 £3 £4 Profit return per £1 £1.87 by month 4 short-term profit ROI £4.11 by month 24 with sustained long-term effects included Short term Sustained long-term effects, months 4 to 24 0 4 8 12 16 20 24 Months after investment
Reported profit ROI more than doubles once long-term effects are counted. A 30-day measurement window sees only the leftmost sliver of this curve.

The same analysis reports Ipsos MMA work recommending that 50% to 60% of marketing spend go to brand building and 40% to 50% to performance tactics, and a Nielsen study for Google finding that a 1% increase in brand awareness drives a 0.4% increase in short-term sales and a 0.6% increase in long-term sales. Awareness, in other words, pays out more in the long run than in the quarter.

The same conclusion arrives from a different direction in Les Binet and Peter Field's effectiveness research for the IPA, drawn from the IPA Effectiveness Databank across Marketing in the Era of Accountability, The Long and the Short of It, Media in Focus and Effectiveness in Context. That body of work is the origin of the 60:40 split between long-term brand building and short-term sales activation, and their more recent argument is that the industry has drifted towards efficiency, targeting and short-term metrics at the expense of scale, reach and genuine brand-building effectiveness.

Setting the measurement window

Adstock and carryover are not modelling exotica. They are the reason a four-week read on a brand campaign will always look worse than the campaign was.

  • Set the window from your own data. Look at the observed lag between exposure and conversion in your closed-won or repeat-purchase records, not at a reporting convention.
  • State the window on every effectiveness number you publish. "ROMI of 10% over an eight-week window" is a claim. "ROMI of 10%" is not.
  • Where the window has to be short, say what it excludes. Reporting a short-window figure as the total return is the most common way effectiveness gets understated for brand activity and overstated for retargeting.

What to do when the numbers disagree

This is the most common real-world situation and the one almost no guide addresses. The platform says a 4x return, the model says 1.6x, the lift test says 2.1x. All three can be computed correctly and still disagree, because they answer different questions over different windows against different counterfactuals.

A workable order of precedence:

  1. A well-powered experiment wins. It has a real control group, so it is the closest thing to ground truth available for the channel and window it covered.
  2. A calibrated model beats an uncalibrated one. If the MMM has been fitted with experimental priors, its estimates carry that evidence forward into periods and channels the test did not cover.
  3. Platform-reported numbers rank last for contribution. They remain the right tool for in-flight pacing, creative rotation and bid management, where relative signal matters more than absolute truth.

Before reconciling anything, check the mundane explanations. Mismatched date ranges, different attribution windows, revenue gross versus net of returns, one source counting orders and another counting sessions. In practice these account for a large share of apparent contradictions.

Then treat the residual disagreement as information rather than noise. A channel where the platform claims far more than both the model and the test find is a channel harvesting demand that already existed. That is a finding, and it is usually actionable.

A marketing effectiveness framework you can run

Frameworks fail when they start with tooling. Start with the decision instead.

  1. Name the decision. Write down what will change depending on the answer: a budget shift, a channel cut, a flight extension. If nothing changes either way, do not run the measurement.
  2. Define the outcome and the counterfactual. Which revenue line, at what grain, compared against what baseline. This is the step most teams skip and most disputes trace back to.
  3. Set the window against the sales cycle. Use observed lag, and record the window alongside the result.
  4. Unify the data at the grain the method needs. Weekly by geography for MMM, market-level for geo tests, user-level for MTA. Spend must be working media, net of fees, joined to revenue on a common calendar.
  5. Choose the method your data can support. Use the decision tree above and accept the constraint rather than forcing a method the data cannot carry.
  6. Report ranges, not points. Publish the interval and the assumptions, particularly the margin assumption inside any ROMI figure. Precision you do not have erodes trust at board level faster than uncertainty you disclose.
  7. Bind the result to a decision and record it. State what you changed, when, and what you expect to observe as a result.

Measurement is not the deliverable

The report is not the point. A marketing effectiveness programme that produces a quarterly deck and no reallocation has cost money and returned nothing.

Every effectiveness result should terminate in one of four decisions: increase spend where the incremental return clears the hurdle, decrease it where it does not, shift the mix between brand and activation, or run a further test where the interval is too wide to act on. "Continue monitoring" is not one of them.

Getting from a model output to a defensible allocation is a distinct problem from fitting the model, which is why we treat the interpretation and allocation layers separately in MMM Singularity and Aryma Nebula rather than folding them into the statistical core. The broader argument for that separation is set out in Peripheral Agentic MMM: automate the work around the model, keep the causal core human-led. The allocation those layers produce is the input the next media plan is built on.

The learnings most organisations lose

There is a failure mode underneath every method described above, and it has nothing to do with statistics.

Organisations forget what they already learned. The saturation point found for paid social two years ago sits in a deck on a shared drive. The reason a geo test was designed with 40 markets rather than 20 is in a thread with someone who has since left. The margin assumption behind last year's ROMI figure is in a spreadsheet nobody can locate. Two cycles later, a new team runs the same test, reaches the same conclusion, and pays for it twice.

That is the real cost of treating marketing effectiveness as a series of projects rather than as a system. Each cycle restarts from close to zero, and the institutional understanding that should compound instead resets.

This is the problem MMM Synapse was built for. It is a memory layer over your existing MMM decks, model outputs and reports, turning them into an encrypted, queryable knowledge base, so a question like "what did we conclude about TV saturation across the last three studies" returns an answer rather than a search. It sits alongside MMMGPT, which answers general methodology questions from a decade of MMM knowledge, where Synapse answers questions about your own history.

Adding an intelligence layer to marketing measurement does not mean handing the model to an agent. It means making sure the rigour you have already paid for is still available the next time someone asks the question. Effectiveness compounds only if the learnings do, which is why we treat MMM Synapse as infrastructure for the measurement system rather than an accessory to it. It starts at $99/mo.

Frequently asked questions

What is the difference between marketing effectiveness and marketing efficiency?

Effectiveness asks whether marketing produced the business outcome it was funded to produce, measured in incremental revenue, profit and brand equity over quarters and years. Efficiency asks what each unit of output cost, measured in CPA, CPL and CPM over days and weeks. A campaign can be highly efficient and completely ineffective if it simply reaches people who would have bought anyway.

How do you calculate marketing ROI, and how is ROMI different?

Marketing ROI as commonly reported divides attributed revenue by media spend, which ignores both margin and the fact that some of that revenue would have occurred regardless. ROMI is stricter: incremental gross profit minus media spend, divided by media spend. It answers the question a CFO is asking. Always state the margin assumption inside it, because the result is highly sensitive to it.

What is the difference between MMM and multi-touch attribution?

Marketing mix modelling is top-down. It uses aggregated time-series data to estimate each channel's contribution, including offline media, and models saturation and carryover. Multi-touch attribution is bottom-up, distributing credit across observed digital touchpoints at user level. MMM is causal in intent and survives signal loss. MTA is correlational, cannot see offline, and degrades as consent and cross-device gaps widen.

Does marketing mix modelling work without cookies?

Yes. MMM uses aggregated weekly spend, sales and control variables rather than user-level tracking, so consent requirements, cookie deprecation and device-level privacy changes do not break it. That resilience is much of why MMM returned to prominence. Its real constraints are different: it needs several years of history and genuine variation in spend to identify channel effects reliably.

How long does it take to see results from marketing?

Longer than most reporting windows allow. Google and WARC analysis found that returns in the first four months equal the returns generated across the following 20 months, with profit ROI rising from £1.87 per £1 in the short term to £4.11 once sustained effects are included. Set your window from your own observed exposure-to-conversion lag, and always state it alongside the result.

How often should you measure marketing effectiveness?

Match the cadence to the method. Experiments run when a specific decision needs evidence. Mix models are typically refreshed quarterly, because a shorter cycle adds noise rather than information. Platform reporting is read daily or weekly, but only for pacing and optimisation. Reviewing an MMM monthly usually produces churn in the allocation rather than a better allocation.

AL
Aryma Labs
Aryma Labs

Aryma Labs is a marketing mix modeling consultancy founded in 2019. Aryma AI is its Gen AI division, applying agents to the periphery of MMM while keeping the statistical core human-led.

Explore the Aryma AI suite

Gen AI products for marketing mix modeling, built on a human-led statistical core. Explore the suite, or talk to the team.