Incident failure 2 am

Why IT Incident Processes Fail at 2 am (And What Actually Fixes It)

July 02, 20266 min read

Why Your Incident Process Works in a Meeting Room and Fails at 2 am

I've run a lot of incident simulations over the years.

In a well-designed tabletop exercise - a fictional scenario, the right people in the room, a structured facilitator - teams almost always perform well. Roles get filled. Communication templates get used. Escalation decisions get made at the right threshold. The post-exercise debrief is positive. People leave feeling more confident about their incident readiness than when they arrived.

And then a real incident happens at 2 am on a Sunday, and the performance looks nothing like the simulation.

This gap - between how a process works in a controlled environment and how it holds up under real conditions - is one of the most consistent and underappreciated challenges in incident management. I want to explain why it exists and what actually closes it.

What Changes at 2am

The meeting room simulation has a set of characteristics that make performance easier than the real thing.

Everyone is rested. The scenario is known to be fictional, which reduces the emotional load. The facilitator is present to redirect if the team drifts. There are no actual clients affected. The CEO is not actually calling. The stakes are intellectual rather than real.

At 2 am on a Sunday, every one of those conditions is reversed.

The person who takes the first call is just waking up. Their cognitive function is not at its peak. The information available to them is fragmentary and arriving in an unstructured way. The situation is real, which means the emotional load is real. The pressure from above - from leadership, from clients, from the awareness that this is costing money every minute - is real.

And the process that was practised in the meeting room is now competing with something much more powerful: instinct, habit, and the desire to just fix the problem as fast as possible without the overhead of following a structure.

Why Instinct Overrides Process Under Pressure

There is a well-established principle in performance psychology - relevant to everyone from surgeons to pilots to incident commanders - that under conditions of high stress and cognitive load, people revert to their most deeply ingrained habits.

Not to their most recently learned behaviours. To their most deeply ingrained ones.

This means that a process practised twice in a workshop will lose to a habit formed over years of improvised incident response, every time, under real pressure. The engineer who has handled incidents informally for three years will do those things at 2 am regardless of what the playbook says. Not deliberately. Because the habit is stronger than the training.

This is not a character flaw. It is a predictable response to stress that applies to everyone. The surgical checklist movement - which dramatically reduced surgical errors - was built on exactly this insight: that even the most highly trained professionals, under stress and cognitive load, miss steps they know perfectly well. The checklist does not work because they don't know the steps, but because it overrides the tendency to skip under pressure.

The implication for incident management is direct. The process needs to be practised enough - in conditions close enough to the real thing - that it becomes a deeply ingrained habit. Not the improvisation.

What "Close Enough to the Real Thing" Actually Means

A tabletop exercise where everyone is comfortable, the scenario is obviously fictional, and the facilitator rescues the team when they get stuck builds familiarity with the process. That's useful. But it doesn't build the specific capability that matters at 2 am.

Building that capability requires practice conditions that introduce at least some of the friction of the real situation. Time pressure. Incomplete information. Decisions that have to be made without certainty. Communication that has to be drafted quickly, not considered at leisure.

It also requires frequency. A quarterly tabletop is significantly better than nothing. A monthly practice - even informal, even short - builds the kind of muscle memory that holds up under real conditions. The teams I've seen handle major incidents most effectively are the ones where the process has been rehearsed often enough that it feels automatic rather than effortful.

The 2 am Escalation Decision Problem

There is a specific moment in every late-night incident that determines whether the process is followed or bypassed: the moment when the on-call engineer must decide whether to escalate.

Escalation means waking people up. It means taking on the social and professional responsibility of having judged the situation serious enough to require it. If that judgment turns out to be wrong, there is a social cost.

This social cost is rarely explicit. But the culture around escalation - whether it is genuinely safe to escalate and be wrong - determines how the on-call engineer makes that 2 am decision.

In organisations where the culture does not support early escalation, engineers wait until they're certain before escalating. By the time they're certain, thirty minutes have passed. The first client communication is late.

The escalation threshold needs to be defined precisely enough to be applied without interpretation under pressure. And the culture needs to make it genuinely safe to escalate at that threshold - which means leaders responding to an escalation that turns out to be unnecessary with appreciation rather than irritation.

The Documentation That Needs to Survive 2 am

Critical incident documentation needs to be accessible within 30 seconds on a mobile device at 2 am, without requiring a VPN connection or a login to a system the on-call engineer hasn't opened in 2 months.

If the on-call engineer is reaching for the escalation path from their phone at 2 am and it takes more than a minute to find it, they will not use it. They will do what they remember. And what they remember is the habit, not the process.

The documentation should be simple, mobile-accessible, and pinned somewhere that requires no navigation to reach. A pinned message in the incident Slack channel. A note in the on-call calendar event. Simple enough to sound almost trivial - important enough to determine whether the process gets used when it matters most.

Frequently Asked Questions

Why does incident response performance degrade at night or on weekends?

Cognitive performance degrades with fatigue, which means engineers responding at 2 am are working with reduced capacity for complex decision-making. Combined with the emotional pressure of a real incident versus a simulation, this leads to a reversion to deeply ingrained habits rather than to recently learned process steps.

How can I help my on-call team feel more confident during real incidents?

Frequency of practice matters more than the sophistication of that practice. Regular, short exercises - even informal walkthroughs - build the kind of muscle memory that holds up under real conditions. The goal is to make the process feel automatic rather than effortful when the pressure is high.

What is the best format for on-call incident documentation?

Documentation needs to be mobile-accessible without requiring VPN or complex logins. The most effective formats are pinned messages in the team's incident Slack channel, notes in on-call calendar events, or a simple shared document with a short, memorable URL. The test: can any member of the on-call rotation reach the escalation path from their phone in under 30 seconds at 2 am?

How often should an incident response process be practised?

At a minimum, quarterly, with brief informal rehearsals monthly where possible. Research and practical experience consistently show that the gap between simulation performance and real-world performance narrows as the frequency of practice increases.

Back to Blog