What an Email Programme Actually Returns
Published 12/2/2025 · 12 min read · Marketing & SEO tools
The email ROI figures everyone quotes come from self-reported surveys. Litmus's State of Email puts the return around 36 to 1 across roughly five hundred marketing professionals, with about a fifth of respondents saying they were unsure of their own ROI; the DMA's UK tracker asks around 250 marketers and lands in the same neighbourhood. These are real, named sources, but they measure what marketers believe, and belief here has two systematic errors. The denominator usually contains the platform fee and not the people; the numerator usually contains revenue that would have arrived anyway. Fix both and the picture changes. Take a 500,000-person audience with a 10% holdout: treated subscribers spend $4.20 each and the holdout spends $3.10, so incremental revenue is $1.10 per person, or $495,000 across the treated base. Last-click attribution credited email with $1,150,000 — 2.32 times as much. Full cost including 1.5 fully loaded people is $260,000. The same programme therefore reports 47.92 to 1 on the usual arithmetic, 1.90 to 1 on incremental revenue against full cost, and 0.86 to 1 on incremental gross profit. Only a holdout can tell you which one you have.
The famous email ROI multiples are self-reported survey figures with the labour left out of the denominator and revenue that would have arrived anyway left in the numerator. Build the honest version instead: incremental revenue against full cost, measured with a holdout — and the arithmetic of how big that holdout has to be.
Where the famous multiple comes from
The figure is not invented, and it is not a secret. Litmus's State of Email survey reports an email return in the region of 36 to 1, drawn from roughly five hundred marketing professionals; the Data & Marketing Association's UK Marketer Email Tracker asks around 250 marketers and reports something of the same order. Both name their sample and their method. The method is the problem: they ask marketers what their return is, and they publish the answers.
Litmus is candid about a detail that ought to travel with the number and never does: about a fifth of respondents said they were unsure of their current ROI. A survey in which a fifth of the sample does not know the answer is still useful for describing sentiment. It is not a measurement of a channel. And the four fifths who did answer were not audited; they reported whatever their attribution tool told them, which brings us to the two systematic errors.
The denominator: what an email programme costs when you count the people
The reason email's reported return is so much higher than any other channel's is that its cost is usually taken to be a software subscription. Everything else — the person who writes the campaigns, the designer who builds the templates, the developer who maintains the integration, the agency retainer, the money spent acquiring subscribers in the first place — sits in a different budget line and never reaches the calculation.
Cost the programme honestly for one year: $24,000 for the sending platform, $135,000 for 1.5 fully loaded people, $30,000 of creative and agency work, $60,000 spent acquiring subscribers, and $11,000 of deliverability tooling and data. That is $260,000 — more than ten times the platform fee, and the platform fee is the only line most calculations contain. The subscriber acquisition line is the one people argue about, and it belongs there: a list is not free inventory, it is an asset somebody paid for, and if you exclude its cost you are valuing a channel by pretending its input arrived by magic.
The numerator: measured revenue is not incremental revenue
Every attribution model answers the question "which touchpoint preceded the purchase". No attribution model answers the question "would the purchase have happened anyway". Those are different questions, and only the second one is about return. For email specifically the gap is wide, because email reaches people who already know you, already bought once, and are already on their way back — which is precisely the population most likely to buy without being emailed.
Last click is the worst offender here, and specifically the worst for email. A person decides to buy, then goes looking for the product, finds your message sitting in their inbox because it is the most convenient link they own, clicks it, and buys. Every step of that story is real, and every step of it would have happened without the message. Last click reads it as a conversion caused by email. So does last non-direct click, which is the default in most analytics tools.
The window setting compounds it. Suppose conversions land 62% on the day of the click, 18% on days one and two, 11% on days three to seven and 9% on days eight to thirty. A one-day attribution window measures $920,000 of email revenue, a seven-day window $1,046,500, and a thirty-day window $1,150,000. Nothing about the business changed between those three numbers; a 25.0% swing came entirely from a dropdown in a reporting tool.
The holdout, and what it reveals
A holdout is a randomly selected slice of the audience that is excluded from the programme. Because assignment is random, the two groups differ only in whether they received the emails, so the difference in their spending is caused by the emails. That is the whole of causal inference in one sentence, and it is the only instrument that answers the question ROI is asking.
Work it through. A 500,000-person audience, a 10% holdout, so 450,000 treated and 50,000 not, over one year. Treated subscribers spend $4.20 each, for $1,890,000. The holdout spends $3.10 each, for $155,000 — note that it is not zero, and that is the entire point. Incremental revenue is $1.10 per person, or $495,000 across the treated base. Meanwhile last-click attribution credited email with $1,150,000. The incremental figure is 43.0% of that, so last click overstated the programme's contribution by 132.3%.
Now divide three ways with the same programme. Last-click revenue over the platform fee alone gives 47.92 to 1, which is how a figure in the range of the survey numbers gets produced. Incremental revenue over full cost gives 1.90 to 1. And because revenue is not profit, apply a 45% gross margin: incremental gross profit is $222,750 against $260,000 of cost, a ratio of 0.86 to 1 and a net contribution of minus $37,250. One programme, three defensible arithmetics, and only the third one tells you whether to keep funding it.
How big does the holdout have to be?
Big enough that the difference you care about is larger than the noise, and no bigger, because every person in the holdout is revenue you chose not to earn. Revenue per person is heavily skewed — most subscribers spend nothing and a few spend a lot — so the standard deviation is large relative to the mean. Take it as $18.00 per person against a mean of $4.20, and the standard error of the difference is the standard deviation times the square root of one over each group size added together.
On the 500,000-person audience with a 10% holdout, that standard error is $0.0849 per person. The observed gap of $1.10 is therefore about thirteen standard errors — overwhelming, and a sign the holdout is larger than it needs to be. The smallest effect this design can detect at 80% power and a 5% two-sided significance level is $0.238 per person, which is 5.66% of baseline spending and $106,975 of annual incremental revenue across the treated base. Running that holdout for a year costs $55,000 of forgone revenue.
Shrink it and the trade becomes visible. A 2% holdout on the same list forgoes only $11,000 but can only detect $0.509 per person, so it would still have caught this programme's effect comfortably while costing a fifth as much. Go the other way — a 50,000-person list with a 10% holdout — and the detectable effect rises to $0.752 per person, 17.9% of baseline, which means a small list can only ever measure a large effect. That is not a flaw in the method; it is the honest statement that small programmes cannot measure themselves precisely and should run the holdout for longer instead of not at all.
Running it without breaking anything
Randomise on a stable identifier so a subscriber does not drift between arms mid-year, and exclude transactional and legally required messages from the holdout — you are testing the marketing programme, not withholding an order confirmation. Keep the arm assignment out of the send tool's segment logic and in the customer record, because a segment built on behaviour will silently reassign people who behave differently, which is exactly the population you are trying to compare.
Measure over a window long enough to contain the purchase cycle and stop looking at the result before it ends — checking a holdout weekly and calling it when the gap first looks significant is the same error as stopping an A/B test early, and the companion article on test sizing in this series explains what that does to a false-positive rate. Report the incremental figure alongside the attributed one rather than replacing it, because the attributed number is what everyone else in the organisation is looking at, and a metric that appears without its predecessor gets argued with rather than used.
| Audience size | Holdout share | Smallest detectable effect per person | As a share of baseline spending | Revenue forgone in the holdout, per year |
|---|---|---|---|---|
| 500,000 | 10% | $0.238 | 5.66% | $55,000 |
| 500,000 | 2% | $0.509 | 12.13% | $11,000 |
| 50,000 | 10% | $0.752 | 17.90% | $5,500 |
| 50,000 | 20% | $0.564 | 13.42% | $11,000 |
Worked with our own calculator
Email marketing ROI calculator
Given
- Emails sent
- 5,000
- Open rate (%)
- 27
- Click rate of opens (%)
- 13.5
- Conversion rate of clicks (%)
- 3.6
- Average order value
- $38.00
- Campaign cost
- $250.00
Result
- Orders
- 7
- Revenue
- $249.32
- Net profit
- -$0.68
- ROI
- -0.27%
These figures are produced by the calculator below, not typed in by hand — they are recomputed whenever the tool changes.
Run it on your own figures →Frequently asked questions
- Is the widely quoted email ROI figure wrong?
- It is not fabricated, it is just measuring something else. Litmus's State of Email puts the return around 36 to 1 from a survey of roughly five hundred marketing professionals, and about a fifth of them said they were unsure of their own ROI; the DMA's UK tracker asks about 250 marketers and reports a similar order of magnitude. Both are self-reported, unaudited, and computed by respondents using whatever attribution and whatever cost base they happen to use. Treat them as a measure of what practitioners believe, which is genuinely interesting, and not as a benchmark your own programme should hit.
- What belongs in the cost side of an email ROI calculation?
- Everything you would stop paying for if you shut the programme down. In the worked example that is $24,000 of sending platform, $135,000 for 1.5 fully loaded people, $30,000 of creative and agency work, $60,000 spent acquiring subscribers and $11,000 of deliverability tooling and data — $260,000 in total, against a platform fee of $24,000. The people line is the one usually missing and the largest single item. The subscriber acquisition line is the one usually argued about: a list is an asset somebody paid for, and leaving its cost out means valuing the channel as though its input were free.
- How large should the holdout be, and what does it cost me?
- Size it from the smallest effect you would act on, not from a round percentage. With revenue per person averaging $4.20 at a standard deviation of $18.00, a 10% holdout on a 500,000-person list detects effects down to $0.238 per person — 5.66% of baseline — and forgoes $55,000 of revenue a year. A 2% holdout on the same list detects down to $0.509 and forgoes $11,000. On a 50,000-person list a 10% holdout only reaches $0.752, which is 17.9% of baseline, so smaller programmes need either a larger share or a longer window. The cost is real and computable, which is what makes it a decision rather than an argument.
- Why is last-click attribution especially generous to email?
- Because email is very often the last convenient link before a purchase that was already decided. Someone intends to buy, looks for the product, and the message in their inbox is the shortest path to it; they click, they buy, and the model records a conversion caused by email. In the worked programme last click credited $1,150,000 while the holdout showed $495,000 of incremental revenue, so 57.0% of the credited revenue would have arrived anyway. The attribution window makes it worse: on the stipulated conversion lag, a one-day window reports $920,000 and a thirty-day window $1,150,000 — a 25.0% difference produced by a setting.
- Is an email programme still worth running if the honest ratio is close to 1?
- It depends what the arithmetic is telling you to change, not whether to stop. A programme returning 1.90 to 1 on incremental revenue and 0.86 to 1 on incremental gross profit is not a channel to abandon; it is a channel whose cost base is too heavy for what it currently produces, and the levers are visible in the same calculation. Cut send frequency to the segment that is not responding, move the fixed creative cost into templates that get reused, or raise the incremental effect by targeting the people the holdout shows are actually being moved. Then re-measure. The point of the honest ratio is that it makes those choices computable, which the 36-to-1 figure never did.
Articles you may find interesting
All guides →Related tools
Sources
- Litmus — State of Email — self-reported email ROI from a survey of marketing professionals, including the share of respondents unsure of their own return
- Data & Marketing Association (UK) — Marketer Email Tracker — annual survey of UK marketers reporting email return on investment
- Google Analytics Help — Attribution models and lookback windows — how last-click and last non-direct click assign credit
- Harvard Business Review — Blake, Nosko & Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness" — the incrementality gap between attributed and causal advertising effects
- Journal of Marketing Research — Gordon, Zettelmeyer, Bhargava & Chapsky — comparing observational attribution with randomised experiments
Spotted a mistake in this article?