Bayesian estimation gets talked about as if it were a more advanced, more accurate kind of Marketing Mix Modelling (MMM). Plenty of practitioners learned ordinary least squares (OLS) and aren’t quite sure what changes when a model goes Bayesian. Plenty of clients are told it’s better without being told why.
This is an explainer, not an argument. I have no stake in either method: both are useful, and both can be used badly. The aim is to show how OLS estimation works compared with Bayesian estimation, using one small example worked by hand – and what that means when you’re reading or buying a model. Which method suits which situation is a question for another piece.
The example
Eight weeks of shampoo sales. Each week has a price per bottle and the volume sold, in thousands.
The example is deliberately kept small, so that every step can be worked by hand. No modeller would build a real model on eight data points: a real Marketing Mix Model would typically have a hundred or more – three years of weekly data gives 156. The logic is the same; only the arithmetic gets longer.
| Week | Price (£) | Volume (000) |
|---|---|---|
| 1 | 2.50 | 118 |
| 2 | 2.80 | 104 |
| 3 | 2.30 | 126 |
| 4 | 3.00 | 95 |
| 5 | 2.60 | 113 |
| 6 | 3.20 | 92 |
| 7 | 2.40 | 119 |
| 8 | 2.90 | 101 |
When the price is high, volume tends to be low. The question is by how much: how many fewer bottles for every £1 on the price? That number is the price slope, and both methods are trying to estimate it.
OLS: the line that misses by the least
OLS draws the straight line that misses the points by the least, in total. It’s the exercise most of us did at school – a chart full of dots and a ruler – done by the numbers instead of by eye. Measure each miss, up or down, from the line. Square each one, so the ups and downs can’t cancel. Add the squares up. The best line is the one with the smallest total.
You can do it by hand in three steps:
- Averages and gaps. The average price is £2.7125 and the average volume is 108.5 thousand. For each week, take the gap between its price and the average price, and the same for volume.
- Multiply the gaps. For each week, multiply the price gap by the volume gap, and square the price gap. Across the eight weeks, the products add to −26.75 and the squares to 0.6888.
- Slope and starting level. The slope is −26.75 ÷ 0.6888 = −38.84. The starting level is the average volume minus the slope times the average price: 108.5 − (−38.84 × 2.7125) = 213.85.
| Week | Price gap | Volume gap | Gap × gap | Price gap squared |
|---|---|---|---|---|
| 1 | −0.2125 | 9.5 | −2.02 | 0.0452 |
| 2 | 0.0875 | −4.5 | −0.39 | 0.0077 |
| 3 | −0.4125 | 17.5 | −7.22 | 0.1702 |
| 4 | 0.2875 | −13.5 | −3.88 | 0.0827 |
| 5 | −0.1125 | 4.5 | −0.51 | 0.0127 |
| 6 | 0.4875 | −16.5 | −8.04 | 0.2377 |
| 7 | −0.3125 | 10.5 | −3.28 | 0.0977 |
| 8 | 0.1875 | −7.5 | −1.41 | 0.0352 |
| Total | −26.75 | 0.6888 |
So the fitted line is volume = 213.85 − 38.84 × price. Each 10p on the price lowers volume by about 3,900 bottles a week.
How sure is OLS?
The points don’t sit exactly on the line. Square the misses and they add to 19.07. Divide by six – eight weeks, minus the two numbers we estimated – and take the square root: the typical miss is 1.78 thousand units.
From that comes the standard error of the slope: how far it would wobble if we had a different eight weeks. It’s the typical miss divided by the square root of 0.6888, which gives 2.15.
A 95% confidence range is the slope plus or minus a multiplier times the standard error. With a large sample the multiplier is 1.96. With only eight data points the standard error is itself uncertain, so the multiplier comes from a t distribution with six degrees of freedom – a normal curve with fatter tails – and is 2.447. That gives ±5.26: a range of −44.1 to −33.6. More data shrinks the multiplier: 2.23 with ten degrees of freedom, 2.04 with thirty, and close to 1.96 with the hundred or more data points a real model would have.
The 95% describes the method rather than this particular range: build a range this way from many different samples, and 95% of those ranges would contain the true slope. And because this range sits well clear of zero, we can be confident that a higher price really does reduce volume – the slope is, in the usual phrase, statistically significant.
Bayesian: start with a belief
A Bayesian estimate starts somewhere else: with what we believe before seeing the data. For this example, assume earlier shampoo studies suggest a price slope of around −15, with a spread of 5. That’s the belief.
The eight weeks are the evidence. On their own, as OLS showed, they point to −38.84 with a standard error of 2.15. The Bayesian answer is a blend of the two, leaning towards whichever source is more certain.
The blend
Certainty here has a precise meaning. Each source’s precision is one divided by its spread squared:
- Belief: 1 ÷ 5² = 0.040
- Evidence: 1 ÷ 2.15² = 0.217
The combined precision is 0.257. Each source’s weight is its share of that: 84% for the evidence and 16% for the belief. The blended slope is the weighted average:
(0.217 × −38.84 + 0.040 × −15) ÷ 0.257 = −35.12
Its spread is 1 ÷ √0.257 = 1.97, so the 95% range is −39.0 to −31.3. In this simple version, the typical miss from OLS is treated as a known noise level.
The result sits between the belief and the evidence, close to the evidence, because the evidence is more precise.
How much the belief matters
Change the belief’s spread and the answer moves:
| Belief spread (around −15) | Share from the data | Slope | Spread |
|---|---|---|---|
| 1 | 18% | −19.25 | 0.91 |
| 2 | 46% | −26.07 | 1.46 |
| 5 | 84% | −35.12 | 1.97 |
| 10 | 96% | −37.79 | 2.10 |
| 50 | 100% | −38.79 | 2.15 |
| No belief (OLS) | 100% | −38.84 | 2.15 |
A tight belief pulls the answer towards −15: with a spread of 1, only 18% of the answer comes from the data. A vague belief leaves the answer where OLS puts it. OLS isn’t a different world. It’s what a Bayesian estimate gives you when the belief carries no weight.
Side by side
| OLS | Bayesian | |
|---|---|---|
| Price slope | −38.84 | −35.12 |
| 95% range | −44.1 to −33.6 (confidence) | −39.0 to −31.3 (credible) |
| A 10p price rise | −3.88 thousand units | −3.51 thousand units |
| What the range means | Built so that 95% of ranges made this way would catch the true slope | Given the data and the belief, a 95% chance the slope lies inside |
| Starting belief needed | No | Yes |
Three things to keep in mind
The narrower range isn’t free. The Bayesian range is narrower partly because it borrows certainty from the belief, and partly because of simplifications in this example: the noise level is treated as known, and the multiplier is 1.96 rather than 2.447. Certainty borrowed from a belief is only as good as the belief.
A confident belief can be confidently wrong. In this example the belief was a long way from the data, and the data was clean enough to win. In a real Marketing Mix Model, many channels have thin or tangled data – spend that barely varies, or moves in step with other channels – and there the belief can do most of the work while the result still looks precise.
Real models aren’t done by hand. A Marketing Mix Model estimates dozens of numbers at once, and a Bayesian one usually gets there by computer simulation rather than a formula. But the logic is the same: belief and evidence, each weighted by how certain it is.
What each method gives you
OLS uses only the data. It’s quick to calculate and easy to check. With noisy data, or little variation in what it is measuring, it can swing to an implausible number, and it has no way to use what is already known.
Bayesian estimation can use that knowledge. It shows where the answer comes from – a weighted average of belief and evidence, with weights that follow precision – and it makes the assumption visible, so it can be tested. Its range is a direct statement: given the data and the belief, the slope is in this range with a 95% chance.
If you’re told Bayesian is better
Ask four questions:
- What did you believe before you saw the data, for which numbers, and where did those beliefs come from?
- How much of each result came from the data, and how much from the belief?
- What happens to the results if the beliefs are loosened?
- Where the data and the beliefs disagreed, which won – and why?
A good modeller will answer all four without hesitation. Neither method is a mark of quality on its own. How it’s used is.
If you’d like a second opinion on how your model was estimated, drop a line to [email protected].