Calibrated Confidence: Why Framing Decisions as Bets Exposes Sloppy Thinking

You can be sure and wrong at the same time, and it feels identical to being sure and right. The only way to tell the two apart is to put a number on the call and keep score, which is what turns a feature bet into something you can actually learn from.

8 min read · for the tool Bet Framing

You’re greenlighting a build. The team wants two quarters and four engineers to ship a feature you’re sure the market wants, and you say it in the room: “I’m confident this lands.” Everyone nods, the budget moves, the work starts. Six months later it ships to a shrug. Usage never crosses the line you had in your head, and the post-mortem produces the usual phrases. Market wasn’t ready. Timing was off. The competitor moved first.

Here’s what you can’t recover from that meeting. You never wrote down how sure you actually were, so now you can’t tell whether this was a reasonable bet that lost or a bad bet you’d make again tomorrow. “Confident” covered both, which is exactly why it was safe to say and useless to learn from. And the next build you greenlight, you’ll walk in feeling the same way, with no record of whether that feeling has ever been worth anything.

The evidence

What you experience as a single sensation is really two things rolled together: how sure you feel, and how often you turn out right. Those two numbers come apart, and the gap between them has been measured across decades of forecasting work. When people say they’re 90% sure, they tend to be right closer to 70 or 75% of the time. When they say it’s a coin flip, they’re often right 60 to 65% of the time. The drift runs in a consistent direction, it’s sizeable, and it is completely invisible from the inside. This is calibration, the match between your stated confidence and how things actually land, and the headline finding is that most people, experts included, are badly calibrated and have no idea.

You can’t feel the gap because overconfidence feels exactly like well-founded confidence. There’s no internal alarm that distinguishes “sure and right” from “sure and wrong.” From where you sit, greenlighting the feature that flops and greenlighting the one that takes off feel the same in the moment of deciding. The only way to separate them is from the outside, by writing the number down and checking it against what happened.

That’s where the second finding bites. The feedback that should fix your calibration gets corrupted before it reaches you, because your memory rewrites your old confidence to fit the result. After a feature underdelivers, people consistently recall having been less sure than they were. After one succeeds, they recall having called it. This is hindsight bias, and it means reflection alone can’t repair your judgement. The record your memory hands you already agrees with the outcome, so it has nothing to teach. You need a number fixed in writing before the result lands, because that’s the one version the result can’t reach back and edit.

The forecasters who beat the field weren’t smarter or better informed. What set them apart was that they stated their beliefs as specific probabilities, kept score, and adjusted when they were wrong. Their 70% calls came true about 70% of the time. Their confidence had been dragged, over many tracked predictions, into line with reality.

How it works

Look at what changes the instant you convert “I’m confident this lands” into “I’m 65% confident this feature clears 10,000 weekly active users by the end of Q3.” The first version is sealed. There is no future fact that can prove it wrong, because “confident” never said what would count as right. The second version is a hypothesis. It names the outcome, the threshold, and the date, and it carries a number you can be held to, which means it can lose, and that is the whole point of it.

Naming 65% forces work that “confident” lets you skip. You’ve just said, out loud, that better than one time in three this thing misses. So you have to look at that third. What does the failure path look like, and what would you see early if you were on it? The number drags the downside into the room while you can still do something about it, instead of leaving it for the post-mortem.

The precision does a second job: it grades your evidence for you. Try to choose between 60% and 80% on this build and notice whether you can. If the two feel indistinguishable, that’s not a rounding problem, it’s a signal that your evidence is too thin to tell a strong bet from a marginal one. Strong forecasters reach for fine-grained numbers, 65% or 72%, because they’ve actually interrogated the case. Weak ones cluster on round, comfortable figures. The number you can defend is a readout of how hard you’ve thought.

A number you wrote down before the result is the one record your memory can’t rewrite to agree with how things turned out, which is why it’s the only part of the decision that can still teach you anything afterward.

How to use it

Before you commit the engineers, write the bet down in one line. “I’m ___% confident that this feature reaches [specific usage threshold] by [specific date].” If you can’t land on a number, you’re not ready to fund it, because you don’t yet know what you’re betting on. That refusal is information too. It sends you back to find the evidence that would let you set a number you’d stand behind.

Then treat the bet as live, not as a one-time pronouncement. As the build moves, the probability should move with it. Early usability tests come back rough and you drop from 65% to 45%. A pilot cohort sticks better than you expected and you climb to 70%. Saying “I was at 65% in March, the beta data put me at 50% in May” is just thinking out loud with a number attached, and it heads off the usual trap of locking onto your first guess and then defending it against everything that arrives after. Moving the number is the bet doing its job, not a sign you’re backing down.

The part that pays the most is the cheapest: keep the written estimate so you can check it later. When Q3 closes, you put the result beside the 65% you actually recorded, not beside the doctored memory of how sure you were. One bet barely tells you anything. Ten or fifteen of them, lined up, start telling you about yourself. Maybe your 70% calls on new features come in nearer 40%, and you’ve found a specific overconfidence in your build decisions that no amount of careful reflection would ever have surfaced. That correction reaches past this one feature into every feature call you make from here.

Why it matters

The room rewards the wrong thing. The leader who says “I’m completely sure this is the right build” reads as decisive and gets the budget. The one who says “I’m 60% on this, and here’s what would move me to 80%” reads as wobbly, even though that second person has told you far more and is far more likely to be right. So people learn to perform certainty regardless of what they actually believe, because the social return on sounding sure beats the return on being accurate.

The cost of that habit is real and almost never shows up on anyone’s ledger. A feature gets full funding and two quarters of headcount when the honest read was 55%, which should have bought a small staged test, not the whole build. Calibrated confidence still leaves you free to act. The forecasters who tracked their numbers weren’t frozen by them. They moved on 70% when 70% was enough, and when it wasn’t, they sized the bet to match. The difference was that they knew it was 70%, and they’d already planned for the 30% that the confident version of them would have waved away.

References

  1. Duke, A. (2018). Thinking in Bets: Making Smarter Decisions When You Don't Have All the Facts. Portfolio.
  2. Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press.
  3. Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown.
  4. Fischhoff, B., & Beyth, R. (1975). I knew it would happen: Remembered probabilities of once-future things. Organizational Behavior and Human Decision Processes, 13(1), 1–16.
  5. Keren, G. (1991). Calibration and probability judgements: Conceptual and methodological issues. Acta Psychologica, 77(3), 217–273.
The newsletter

One tool a week

How you think, decide, lead, focus, and stay steady under pressure. A specific way to practice one move before the next seven days are out. Grounded in evidence, not self-help.

One email a week. Leave whenever. Powered by Buttondown.