Regression to the Mean: Why Extreme Results Lie About What Comes Next

An extreme result will mostly correct itself, yet the change you made just before it gets the credit. The trick is to estimate how much of the result was going to fade anyway, so you stop training yourself on noise.

8 min read · for the tool Regression to the Mean

Your support team has its worst month on record. Response times double, the satisfaction score falls off a cliff, and a couple of accounts threaten to leave. You move fast: a new team lead, a rewritten triage process, a tighter rota. The next month, the numbers recover. The score climbs most of the way back and the accounts stay. The new lead looks like a great hire, and the process rewrite looks like the reason.

Here is the problem you can’t see from inside that story. The recovery would very likely have happened anyway, with or without the new lead and the new process. A month that bad was partly a run of unusual luck, and luck that bad rarely repeats two months running. You changed three things and the number went up, so you’ll keep all three and tell yourself they worked, never finding out which, if any, actually did. This is regression to the mean, and the reason it’s so expensive is not that you misread one month. It’s that you’ve just taught yourself a lesson that isn’t true, and you’ll apply it the next time the stakes are higher.

The evidence

The pattern was first pinned down in the heights of parents and their children. Exceptionally tall parents tended to have children who were tall but closer to average, and exceptionally short parents had children who were short but, again, closer to average. The early reading was biological. It turned out to be something simpler and more general: a property of any measurement that carries a slice of randomness. Whenever a result mixes real signal with chance, an extreme value tends to be followed by a less extreme one, because the chance part is unlikely to land at an extreme twice in a row.

It shows up wherever variable performance gets measured and acted on, and it’s among the most reliably reproduced findings in decision research. One of the cleaner demonstrations comes from flight training. Instructors were convinced that praising a cadet after an excellent flight made the next flight worse, while criticising a poor flight made the next one better. They had the cause exactly backwards. An exceptional flight is partly a lucky flight, so the next one regresses down no matter what you say, and a terrible flight is partly an unlucky one, so the next one regresses up. They were reading pure regression and building a theory of feedback on top of it, one that told them, wrongly, that punishment works and praise doesn’t.

The same effect has been measured in business and in sport. Companies with extreme results drift toward their industry’s normal range over the following years. Players who post a season far above their career norm reliably slide back the next year, and it isn’t complacency or age doing it. It’s the arithmetic of a number that was never going to stay that high. The effect is also one of the most common ways research itself gets fooled: pick a group because their results are extreme, apply anything at all, and the group tends to improve on its own. Without a comparison group held back from the change, the change looks like it worked.

How it works

Treat any result you measure as two things added together. One part is the real, stable level: the genuine quality of your team, the true strength of a product, the underlying close rate of your sales process. The other part is everything that wobbles month to month: timing, who was out sick, which deals slipped, a competitor’s move, plain noise. You only ever see the sum, never the two components separately.

Now think about what it takes to produce an extreme reading. For a result to come in far above or far below normal, the wobble almost certainly pushed hard in the same direction. A record month is usually a good month for the real reasons plus a generous tailwind of luck. The next month, the stable part stays roughly where it was, but the wobble resets to a fresh, independent draw, and a fresh draw is unlikely to be extreme. So the total comes in lower. Nothing about your team changed; the luck simply stopped repeating.

This gives you a way to size the effect instead of just naming it. Two things drive how far a result will regress. The first is how extreme it was: the further from your average, the more of it was probably wobble, and the harder it falls back. The second is how noisy the measure is in the first place. A number that swings wildly month to month is mostly wobble, so almost any extreme reading regresses hard, while a stable, high-signal number barely regresses at all. A record quarter in a lumpy enterprise sales pipeline tells you far less than a record quarter in a steady subscription business.

The more extreme the result and the noisier the measure, the more of it was luck, and the more of it will fade on its own before you’ve done anything at all.

How to use it

The deployable version is simple: when a result is extreme, ask whether it’s a spike or a pattern, and wait for one more data point before you tear anything up. That’s the move to keep in your pocket. Here is the harder version for when you can’t simply wait, because something is on fire and you have to act now.

Before you act, write down where the next result will land if you do nothing. Not a hope, a number. If your team’s normal satisfaction score sits around 80 and this month it cratered to 55, the honest forecast for next month is not 55 and not 80. It’s somewhere in between, pulled back toward 80 by however noisy that score usually is. Say you expect 72 with no intervention. Now you’ve set a baseline that already accounts for the bounce. Make your changes, and judge them against 72, not against the awful 55 that scared you into moving. If next month comes in at 73, your changes did roughly nothing, however much it feels like a turnaround. At 85, you have a real signal worth keeping.

This is what protects you from the worst version of the trap, which is not wasting one decision but corrupting how you learn. Every time you credit a fix for a recovery that was always coming, you write a false rule into your playbook: hire this kind of lead, run this kind of restructure, cut this kind of cost. The rules feel earned because you watched them work. They didn’t work, regression did, and they’ll fail you the day you actually need them to move a number on their own. The flight instructors didn’t just misjudge one cadet. They trained themselves, year after year, to believe in a feedback approach the evidence never supported.

Watch for the moments when this is most expensive: a support metric you’re judging on a single great month, a campaign you’re crediting for one strong quarter, a market you’re about to chase because last quarter spiked. In each case the loud result is partly a fluke, and your job is to forecast the bounce before you build a story on top of it.

Why it matters

Most of the causal stories you tell at work are about extreme results, because those are the ones that grab your attention and demand a response. That’s exactly where regression hides, so a large share of the lessons you’ve drawn from your career may be sitting on a foundation of statistical drift you mistook for cause and effect. The strategy you swear by because it rescued a bad year, the person you promote on type because their first big test went well: some of those judgements are sound. Others are regression wearing a costume, and you can’t tell them apart unless you start forecasting the bounce on purpose.

None of this argues for doing nothing when results go bad. A genuine structural break, a real new competitor, a process that’s actually broken, deserves a fast response, and waiting through it is its own mistake. The skill is telling a signal that warrants action from a number that’s just settling back to where it always sat. You do that by knowing your baseline, expecting the bounce, and refusing to hand the credit to whatever you changed right before the number went the way it was always going to go.

References

  1. Galton, F. (1886). Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263.
  2. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
  3. Secrist, H. (1933). The Triumph of Mediocrity in Business. Bureau of Business Research, Northwestern University.
  4. Schall, T., & Smith, G. (2000). Do baseball players regress toward the mean? The American Statistician, 54(3), 231–235.
  5. Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology, 34(1), 215–220.
The newsletter

One tool a week

How you think, decide, lead, focus, and stay steady under pressure. A specific way to practice one move before the next seven days are out. Grounded in evidence, not self-help.

One email a week. Leave whenever. Powered by Buttondown.