Calibration: The Meta-Skill Behind Every Decision Skill

The probability you say out loud on a go/no-go call sets how hard you commit. If you have never checked those numbers against what actually happened, the figure is probably inflated, and you can find out by how much.

8 min read · for the tool Confidence Calibration

You are in the room where the call gets made. Ship the feature in March or hold it for the next cycle. Someone asks how confident you are that the March build clears testing, and you say eighty percent. The number lands, the room relaxes, and the plan gets built around it. Nobody writes it down. Nobody will ever come back and check whether eighty percent was anywhere close to right.

That stated number did a lot of work. It set how much slack got cut from the schedule, how much budget got held in reserve, how loudly the doubters got overruled. And here is the uncomfortable part: you have no idea whether your eighty percent means eighty percent. It is a feeling that came out of your mouth as a figure, and the figure has never once been tested against what actually happened.

The evidence

When people put a number on how sure they are, the number runs high. This is one of the most reliably reproduced findings in decision research. Pull together the studies and the pattern holds across topics and across people: when someone says they are ninety percent sure, they tend to be right closer to seventy or seventy-five percent of the time. When they say they are certain, dead certain, they are still wrong something like one time in six. The gap is widest on hard questions, which is exactly where you most need the number to be honest.

It is tempting to assume this is a problem for amateurs and that experience burns it off. It does not. When managers are asked to give a range they are ninety percent sure contains the right answer, a range wide enough that they should be wrong only one time in ten, they miss far more often than that. They land the true answer inside their range four to six times out of ten instead of nine. Same overconfidence, and it held steady across industries, across seniority, across how long people had been doing the job. Years in the chair did not fix it.

There is a useful split inside this. Overconfidence comes in more than one flavour, and the one that damages decisions most is overprecision: being too sure your own estimate is tight, too quick to rule out the ways you could be wrong. That is what shows up the moment you say eighty percent on the March build. You are not just claiming the build will probably make it. You are claiming your read on the odds is sharp enough to bet the schedule on, and that is the claim that runs too high.

How it works

The reason this never corrects itself on its own is the feedback loop, or rather the lack of one. Calibration is a skill that needs scoring, and most of professional life never keeps score. In places with fast, clean feedback, weather forecasting, poker, parts of medicine, people get calibrated whether they mean to or not, because the gap between what they said and what happened lands on them quickly and over and over. A weather forecaster who says seventy percent chance of rain finds out by evening, hundreds of times a year.

Your go/no-go calls are the opposite. The feedback comes months late, by which point you have forgotten the exact number you gave. It comes blurred, because the world moved and you can argue the launch was held up by something you could not have known. And it comes already rewritten by hindsight, because once you know the build slipped, you half remember having doubts all along. So the comparison that would tell you whether your eighty percent is any good, the prediction set against the outcome, almost never gets made. The overconfidence just sits there, decade after decade, feeling exactly like well-founded judgement.

A confidence number you never check against outcomes drifts high and stays there, and you have no way of knowing by how much.

This is why calibration sits underneath everything else you do. Every decision tool you reach for ends in a judgement about how much to trust your own read. Run a pre-mortem and you still have to weigh how likely each failure is. Pull a base rate and you still decide how far to override it with what you know about this case. If the confidence feeding all of that runs high, every tool downstream inherits the error. The single number you say out loud on a go/no-go is the cleanest place to see the whole problem, because it is right there in the open, doing damage you can measure.

How to use it

The basic move is to keep score. The next time you make a real go/no-go call, write the date, the call, the number you would say out loud, and the date you will know. Then leave it. When the outcome lands, put it next to what you wrote, not next to the tidied-up version in your head. One score tells you nothing, and five only start to show a shape. By twenty you have a number you can trust.

Once you have a stack, do not just take the average, because the average hides the move that matters. Sort the calls by what kind of call they were. You may find your eighty percent is solid on technical questions, will the build pass, will the system hold, and wildly high on anything involving other people, will they accept, will they deliver, will the client sign. That asymmetry is the usable finding. It tells you precisely where to shave your confidence and where to leave it alone, instead of dragging every number down by the same vague amount.

And here is what to do on the live call before you have a thick record, because the decision will not wait for your sample to fill up. Suppose you have learned that your eighty percent on people-dependent calls actually comes true about sixty percent of the time. The fix is to act on the corrected number rather than trying to feel less sure, because you cannot will a feeling down, and trying just makes you hesitant without making you accurate. When the March question is really a people question and your gut says eighty, you treat it as a sixty out loud and in the plan. You hold more reserve, you keep the fallback warm, you commit at the level the corrected figure justifies. Your gut still reads eighty while the decision runs on sixty, because sixty is what eighty has historically meant for you on calls like this one.

Why it matters

Most of what professional training sharpens is the analysis itself, how to read the market, run the team, judge the build. Almost none of it sharpens the layer above the analysis, the part that decides how much to trust what you just produced. So you end up good at generating a confident read and bad at knowing how much that confidence is worth, which is a strange thing to be expert at and never have checked.

The good news buried in the research is that this is trainable, and not slowly. The people who outforecast professionals with far better access did it mainly by being honest about their own accuracy. When they said seventy percent, things happened about seventy percent of the time. They were not smarter or better briefed. They had simply made the comparison everyone else skips, again and again, until their stated numbers meant what they said.

That is the whole of it. Make the call, say the number, write it down, check it later, and let the gap teach you what your numbers actually mean. The bookkeeping costs minutes. The hard part is being willing to find out that the eighty percent you have been saying for years has been a sixty all along, and then saying sixty next time, out loud, in the room where the call gets made.

References

  1. Lichtenstein, S., Fischhoff, B., & Phillips, L. D. (1982). Calibration of probabilities: The state of the art to 1980. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment Under Uncertainty: Heuristics and Biases (pp. 306–334). Cambridge University Press.
  2. Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517.
  3. Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown.
  4. Keren, G. (1991). Calibration and probability judgements: Conceptual and methodological issues. Acta Psychologica, 77(3), 217–273.
  5. Russo, J. E., & Schoemaker, P. J. H. (1992). Managing overconfidence. Sloan Management Review, 33(2), 7–17.
The newsletter

One tool a week

How you think, decide, lead, focus, and stay steady under pressure. A specific way to practice one move before the next seven days are out. Grounded in evidence, not self-help.

One email a week. Leave whenever. Powered by Buttondown.