One Incident Readiness Question CTO

The One Incident Readiness Question Every CTO Should Be Able to Answer

July 22, 20266 min read

The One Question Every CTO Should Be Able to Answer Before the Next Incident Hits

I've worked with a lot of technically excellent leaders over my 20+ years in IT operations.

People who can diagnose a complex infrastructure failure in minutes. Who have deep knowledge of their systems, their team, and their architecture. Who are respected by their engineers and trusted by their boards. By every conventional measure, highly capable.

And when I ask them one specific question, a significant proportion of them cannot answer it with confidence.

The question is this:

If you were completely unreachable right now - no phone, no laptop, no way to be contacted - could your team handle a major incident end-to-end, including the client communication and the leadership briefing, without you?

Not "would they try." Not "they'd manage somehow." Could they do it well - with structured roles, clear communication, timely client updates, and a confident briefing to the board - without your involvement at any point?

Why This Question Matters More Than Most

Most incident readiness questions focus on the technical response. Can the team diagnose this category of failure? Do they have access to the right systems? Is the runbook current?

Those are important questions. But they address one dimension of incident response - the technical one. And in my experience, the technical response is rarely where the most significant gaps are.

The question above addresses a different dimension: the organisational and communication capability of the team when the person who normally holds it together is not available. It surfaces three separate risks simultaneously.

Single point of failure in leadership. If the honest answer is that the team struggles without you, you are a bottleneck in your own incident response. That bottleneck has a cost accumulating every time an incident occurs. It also has a tail risk: the major incident that happens precisely when you're unavailable - on a flight, in surgery, in a meeting you cannot leave.

Underdeveloped team capability. If the team can't run a P1 without you, it's not because they lack the technical skill. It's because they've never been required to run one without you. The capability is latent. It needs to be developed deliberately - through practice, through permission, through the gradual transfer of responsibility.

Communication and coordination gaps specifically. When I ask this question, the aspect that produces the most hesitation is almost never the technical response. It's the client communication and the leadership briefing. "The team could probably diagnose and fix it" is the common answer. "I'm not sure who would write the client update or how they'd handle the board call" is the honest follow-up.

That's the gap. And it's the gap that, in a live incident without you present, produces the communication vacuum.

The Dependency Is Understandable - And Still a Problem

The dependency that most technically strong leaders create is almost always unintentional and understandable.

You stepped in during incidents because you were the most capable person in the room. You handled the CEO call because you had the relationship and the context. You drafted the client communication because you knew what to say and you knew it needed to be right.

Each of those decisions was individually correct. The cumulative effect is a team that has learned to defer to you, and a set of capabilities - incident command, client communication under pressure, stakeholder briefing - that have not developed in the people who need to have them.

The problem is not that you made those decisions. The problem is that the pattern they created has a cost that is not visible until you're not there.

What It Takes to Be Able to Answer Yes

Transfer the communications lead role explicitly. Identify a specific person - not "someone from the team" - who owns client communication during a P1. Not as a backup to you. As the primary. Give them the templates, the authority to send without your approval, and the direct contact for the key clients they may need to update. Then step back from that role in the next incident and let them do it.

Brief your leadership team on what to expect. Your CEO and board need to understand that the incident briefing will come from someone other than you. That conversation is much easier to have before an incident than during one.

Run an incident simulation without you in the room. Not a tabletop exercise where you're present and can redirect - one where you are explicitly absent and the team has to execute the full response without your input. The discomfort of watching the team do it differently from how you would have done it is not evidence that it went wrong. It is evidence that the capability is transferring.

Ask the question quarterly. Every quarter, ask yourself honestly: if I were unreachable right now, could the team handle this? If the confidence is growing, the work is going in the right direction.

The Answer You're Building Toward

The CTO is on holiday. A P1 hits on a Saturday afternoon. The on-call engineer declares it within five minutes. The incident commander role is filled by a senior engineer who has practised it in simulations. The communications lead sends the first client acknowledgement at the fourteen-minute mark. The CEO receives a structured briefing at twenty minutes. The technical team works the fix. The incident resolves in two hours. The client receives a resolution communication and a follow-up summary within twenty-four hours.

The CTO finds out about it when they land. The team handled it. The client relationship is intact. The post-mortem is already scheduled.

That is what "yes, confidently" looks like. And it is built - deliberately, specifically, over time - through the steps described above. Not through hoping the team will rise to the occasion when the moment comes.

Frequently Asked Questions

What is the single biggest incident management risk for CTOs?

Single-person dependency - the situation where the incident response capability relies on one person being available and engaged. This creates a bottleneck in normal operations and a significant vulnerability when that person is unavailable. The fix is deliberate capability transfer: distributing the communication and coordination skills across the team through practice, not just documenting who the backup is.

How do I know if I'm a single point of failure in my team's incident response?

Ask: who sent the last five client communications during a significant incident? Who briefed the CEO? Who made the escalation decision? If the answer to all three is you, you are a single point of failure. The test isn't whether you have a documented backup - it's whether that backup has actually executed the full response without you.

How do I transfer incident communication ownership to my team?

Start by identifying a specific named person as the communications lead, separate from the technical lead. Give them the templates and the authority to send without your approval. Then deliberately step back from that role in the next real incident. The first time will feel uncomfortable. That discomfort is the capability transfer happening.

Should the CTO be involved in every P1 response?

Ideally not. A well-structured incident process should enable the team to handle a P1 end-to-end without CTO involvement - with the CTO receiving structured updates rather than actively managing the response. The CTO's involvement should be reserved for incidents of genuine strategic severity, not as the default for every P1.

Back to Blog