Measuring content ROI without lying to yourself
Attribution models flatter whoever built them. A measurement framework that survives contact with a sceptical CFO — and tells you something you did not already believe.
The honest problem with content ROI
Content marketing has a measurement problem that is genuinely harder than the equivalent problem in performance media, and pretending otherwise is why so many programmes get defunded in downturns.
Three structural difficulties:
- The returns are lagged and long. A piece published in March may generate its most valuable enquiry in the following January. Quarterly reporting cycles cut straight through that.
- The most influential exposures are the least visible. An article forwarded internally, read in an email client, quoted in a meeting, or summarised by an assistant leaves no trace in your analytics whatsoever.
- The counterfactual is unknowable at the individual level. You cannot know whether this specific buyer would have found you anyway.
None of that means content is unmeasurable. It means content is measurable in aggregate and over time, but not in the way a click-based channel is, and forcing it into that framework produces numbers that are precise and wrong.
The purpose of measurement is not to prove the programme works. It is to find out sooner when part of it does not.Written on the wall of our analytics room
Start from decisions, not metrics
Most dashboards are built by asking what can be measured. Better dashboards are built by asking which decisions get made, then working backwards.
The decisions a content programme actually faces:
- Should we continue investing in this territory, or move the budget?
- Should we produce more of this format, or less?
- Should we update this existing piece, or write something new?
- Should we put paid budget behind this?
- Is this distribution channel worth the ongoing time?
- Should the whole programme be bigger or smaller next year?
For each, ask: what would need to be true to answer it, and what evidence would tell us? Any metric on the dashboard that does not feed one of those decisions is decoration — and worse, it competes for attention with the ones that matter.
The dashboard test
Go through your reporting line by line and, for each number, name the decision it would change and the threshold at which it would change it. Most teams find that between half and two-thirds of what they report fails this test. Removing those numbers makes the remainder considerably more persuasive.
The four layers worth measuring
We measure content in four layers, running from fast and unreliable to slow and meaningful. The mistake is to report only the fastest layer, because it is the one that updates daily.
Layer one: production health
Not outcomes — the machine that makes outcomes possible. Pieces published against plan, average time from brief to publish, proportion needing more than two review rounds, expert hours secured against committed. When outcomes disappoint, this layer usually explains why, and it is the only layer you fully control.
Layer two: reach and engagement
Non-branded organic entrances, returning readers, email click-through, scroll depth on substantial pieces, external citations earned. Useful as leading indicators and as diagnostic detail. Dangerous as headline numbers, because they are easy to inflate and easy to mistake for value.
Layer three: commercial influence
Content-assisted pipeline, self-reported source data, sales-cycle length for content-exposed deals against others, sales usage of specific assets. This is where content connects to money, imperfectly but genuinely.
Layer four: incremental contribution
What would not have happened without the programme. This is the only layer that truly answers the ROI question and the only one requiring deliberate experimental design rather than passive reporting.
Attribution: what it can and cannot tell you
Every attribution model is a set of assumptions about how credit should be assigned, presented as a measurement. It is worth being explicit about what each one systematically distorts:
| Model | Systematically over-credits | Reasonable use |
|---|---|---|
| Last click | Branded search, direct, retargeting | Optimising the final step only |
| First click | Top-of-funnel discovery | Understanding how people arrive |
| Linear | Anything with many touches | A rough sanity check |
| Time decay | Late-stage activity | Short, fast sales cycles |
| Data-driven | Whatever is well-tracked | High volume, single-domain journeys |
| Self-reported | Memorable over frequent | Everything else — genuinely |
The uncomfortable finding, which large-scale studies keep reproducing, is that self-reported attribution — a free-text “how did you hear about us?” on the enquiry form — often correlates better with genuine incremental lift than sophisticated multi-touch models. It has obvious biases. It also captures the podcast, the recommendation from a colleague, and the article read on a phone in an email client, none of which any tracking system will ever see.
Use both. Where they disagree, the disagreement itself is informative: it usually marks a channel doing real work that your tracking cannot observe.
Incrementality tests you can actually run
Incrementality means: what changed because of the thing. It requires a comparison, and content programmes can be compared in several practical ways that do not need a research budget.
Geographic holdout
Where you run paid amplification, exclude a matched set of regions for eight to twelve weeks. Compare enquiry rates. Works well for local and national businesses, less so where your market is thinly spread.
Territory phasing
Launch territories sequentially rather than simultaneously. Each launch becomes a natural experiment against the others. Costs nothing beyond patience and gives you a clean before-and-after per territory.
Pause tests
Stop distributing to one channel for six weeks. Uncomfortable, and the cleanest signal available for ongoing distribution effort. Best run in a quieter trading period so the cost of being wrong is bounded.
Matched-cohort comparison
Compare deals where the buying group had meaningful content exposure against comparable deals where it did not — matched on size, sector and entry point. Not a true experiment, but with enough volume the pattern is informative, particularly for sales-cycle length.
Brand-lift surveys
Two questions to a sample of your target market, quarterly: unaided category association, and whether they have encountered your material. Cheap, and the only instrument that captures influence on people who have not yet identified themselves to you at all.
Modelling the lag
The single biggest reporting failure in content marketing is comparing this quarter's spend against this quarter's return. Content published in a quarter generates returns over the following eight to twelve quarters, and often peaks well after publication.
A workable approach without econometric machinery:
- Take every piece published in a given quarter as a cohort.
- Track that cohort's cumulative contribution — entrances, assisted pipeline, closed revenue — monthly thereafter.
- Plot cohorts on the same axes with months-since-publication on the x-axis.
- After four or five quarters, the shape of the curve becomes clear enough to forecast against.
Two things usually become visible once teams do this. First, the payback period is far longer than anyone assumed — commonly nine to eighteen months in considered-purchase businesses. Second, cohorts differ substantially from one another, and the differences correlate with identifiable changes in editorial approach. That is the finding that improves the programme.
Reporting to people who count money
A finance director's scepticism about content is usually well-founded — they have been shown a lot of impressive-looking numbers that never appeared in a revenue line. The way through is not better-looking dashboards. It is candour about uncertainty.
What earns credibility:
- State the payback period explicitly, with the cohort data supporting it. Long is fine if it is honest and consistent.
- Give ranges, not points. “Between £340k and £610k of influenced pipeline” is more believable and more useful than a single false-precision figure.
- Separate what you measured from what you modelled. Two columns. Always.
- Report what did not work and what you did about it. Nothing builds trust in a number faster than the person presenting it volunteering the failures.
- Compare against the alternative use of the money, not against zero. The real question is never “is content worth it” — it is “is this better than the same money in paid, events, or headcount?”
- Show the compounding asset. Pieces published two years ago still working are the strongest structural argument content has, and it is one paid media cannot make.
Six measurement anti-patterns
- Reporting totals instead of non-branded. Total organic traffic includes people searching your name, who found you some other way. Non-branded is the number that reflects content working.
- Celebrating traffic from irrelevant queries. A viral piece attracting an audience that will never buy is a cost with a nice-looking chart attached.
- Averaging across wildly different pieces. Content performance is heavily skewed — a handful of pieces do most of the work. Medians and distributions tell the truth; means do not.
- Changing the model when the numbers disappoint. If you switch attribution models mid-year, you have lost the ability to compare anything to anything.
- Measuring pieces rather than territories. Individual pieces are noisy. Territories are the unit at which decisions get made.
- Not measuring what you decided not to do. Killed territories and abandoned formats should stay in the record. Otherwise the programme's history reads as an unbroken run of good judgement, which is not a useful thing to believe about yourself.
The last one is the one we have to argue for most often, and it is the one that most improves decision-making over a multi-year programme. A measurement system that only records successes is not a measurement system. It is a marketing asset pointed at your own management.