Systems Over Blame: Why Asking 'What Did the System Allow?' Produces Better Learning Than 'Who Messed Up?'

After an outage, the room wants a name. Finding one feels like progress, but it trains everyone watching to bury their next mistake, which is how the same failure keeps coming back.

8 min read · for the tool Failure Autopsy

A shipment goes out wrong. The customer ordered one configuration and received another, the error reached them before anyone caught it, and now you’re in the review the next morning trying to work out what happened. The pull in the room is immediate and it’s toward a person. Who packed it? Who signed off? Whose initials are on the dispatch note? Somebody is identified, the conversation gets tense, and the meeting lands where these meetings usually land: they’ll be more careful next time, and a line goes into the procedure telling everyone to double-check before dispatch.

Four months on, a different order goes out wrong. Different person on the line, different customer, same shape of mistake. You run the same review, you find the same kind of name, you issue the same kind of reminder. The thing that actually produced both errors, the handoff nobody owned and the check that lived in one person’s head, is still sitting in the workflow exactly where it was, waiting for its next turn.

The evidence

Start with what’s wrong with looking for a name at all. Decades of investigation into the failures that genuinely hurt, in aviation, in hospitals, in power plants, converged on a distinction worth carrying into your own reviews. There’s an old view of human error that treats the mistake as the cause: someone was careless or under-trained, so the failure happened. And there’s a newer view that treats the mistake as a symptom of conditions further upstream. In the new view, “who made the error” is the wrong opening question, because it stops exactly where the useful work begins. The better opening is “what made this error possible, or likely, or close to inevitable?”

The reason that question pays off is structural, and it has a name worth keeping: the Swiss cheese model. Picture every defence you have against a bad shipment as a slice with holes in it. Packing catches most errors but not all. The dispatch check catches some of what packing misses. The customer’s own confirmation catches a few more. A wrong order reaches the customer only when the holes in every slice line up, so a single mistake sails through all of them untouched. The person at the end of that chain isn’t the cause of the failure, just the last slice, the one place where the alignment finally became visible. Blaming them is accurate in the narrow sense and useless in every sense that matters, because the holes were already there before they touched it.

This is where most reviews go wrong, and it’s worth being precise about the cost. If you decide the cause was carelessness, your fix is more care, which depends on a tired human staying vigilant on a Friday afternoon forever. If you decide the cause was an unowned handoff and a check that existed in nobody’s documented process, your fix is to give the handoff an owner and put the check on paper. That fix holds no matter who’s on the line next month, because it takes the slip out of the path instead of asking a person never to slip.

How it works

Here’s the part that makes blame genuinely expensive, beyond just producing a weak fix. When the response to an error is to find and name the person responsible, you’re not only failing to repair the system. You’re teaching everyone who watched the review what happens when an error surfaces, and the lesson they take away is to keep their own buried.

Teams where people feel they can flag a mistake without getting burned report more errors, not because they’re worse but because they’re honest. The errors were always there. What changes is whether you find out about them in a review or find out about them when they hit a customer. The moment a post-mortem turns into a search for someone to carry the blame, the smart move for everyone else in the building is to go quiet: no more flagging the near-miss, no more mentioning the order that almost went out wrong, no more volunteering the workaround they’ve been running because the official process is broken. So you get your one chastened name, and in exchange you lose the early warnings on the next dozen failures, which now arrive with no notice at all.

That’s the real trade a person-centred review makes, and it’s invisible at the time. The cost doesn’t show up in this meeting. It shows up three months later as another wrong shipment that nobody saw coming, because the people who could have seen it coming learned to keep quiet.

The price of naming a person is that everyone else stops telling you where the next failure is forming.

How to use it

So when something has gone wrong and you want the failure to actually teach you something, run the review as an autopsy on the system rather than a search for the guilty. Write three lines. First, what happened, facts only, with every word of blame stripped out: the order shipped with the wrong configuration and reached the customer Tuesday. Second, what the process allowed to go wrong: the handoff from sales to fulfilment had no single owner, and the configuration check lived in one person’s memory rather than in the workflow. Third, the one change you’d make if you ran it tomorrow: the configuration gets confirmed against the order at dispatch by whoever releases the shipment, as a logged step.

Hold yourself to one change on purpose. After a bad outage the urge is to fix everything at once, add ten checks, rewrite the whole runbook, buy a tool. That feels thorough and it mostly buys you process bloat that the team routes around within a month. Find the single condition that contributed most to this specific failure and remove that one. If something similar happens again, the next autopsy surfaces the next-biggest condition, and you peel the system back one real layer at a time instead of burying it under cautions nobody reads.

Give it a day before you sit down. When the wound is fresh, when someone’s still bracing to be blamed, the analysis comes out defensive and you get a confession instead of a diagnosis. A cooling gap of a day or two lets the recrimination burn off so the actual conditions become visible.

The harder case is the one where the person really did slip. They skipped a step they knew, or they were sloppy, and it’s tempting to say the system talk lets them off. It doesn’t. Ask why a single skipped step could reach the customer at all. A well-built process assumes people will occasionally cut corners under pressure and catches it anyway. If one tired decision on a Friday can put a wrong order in front of a client, the missing guardrail is the finding, and “be more careful” is still the fix that changes nothing.

Why it matters

The setbacks that shape an operation are rarely the ones with a dramatic villain. They’re the ones that came back. A delivery error returns under three different names. An outage repeats with a different on-call engineer each time. A handoff drops the same ball for two years, because every review ended at a person and never reached the seam they were standing on. None of those felt like a pattern in the moment. Each one felt like an isolated case of somebody not being careful enough, which is exactly the story that let the real cause survive untouched.

Choosing to ask what the system allowed instead of who messed up isn’t soft, and it isn’t natural. The instinct to find a face for the failure runs deep, and the blame version feels cleaner because it closes the file. The systems version leaves you with a messier answer: several conditions, shared ownership, a process you built that let it happen. That answer is harder to deliver and harder to feel finished with. It’s also the only one that actually keeps the same failure from happening again.

References

  1. Dekker, S. (2006). The Field Guide to Understanding Human Error. Ashgate Publishing.
  2. Reason, J. (1990). Human Error. Cambridge University Press.
  3. Edmondson, A. (1999). Safety to speak up and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.
  4. Perrow, C. (1984). Normal Accidents: Living with High-Risk Technologies. Basic Books.
  5. Senge, P. M. (1990). The Fifth Discipline: The Art and Practice of the Learning Organization. Doubleday.
The newsletter

One tool a week

How you think, decide, lead, focus, and stay steady under pressure. A specific way to practice one move before the next seven days are out. Grounded in evidence, not self-help.

One email a week. Leave whenever. Powered by Buttondown.