
What Does "Incident Ready" Mean for a 50-Person Company? A Specific Checklist
What "Incident-Ready" Actually Looks Like for a 50-Person Company
"Incident ready" is one of those phrases that gets used often and defined rarely.
I've heard it in board presentations, in sales conversations, in post-mortem debriefs. Everyone agrees it's the goal. Very few people have a specific and honest picture of what it actually looks like - what the conditions are that allow you to say, with confidence, that your organisation is genuinely prepared for a major incident rather than hoping one doesn't happen.
I want to give you that picture. Not the enterprise version - the version that is realistic, achievable, and meaningful for a company of 15 to 200 people without a dedicated incident management function.
What Incident Ready Is Not
Let me start by clearing away some common misconceptions.
Incident-ready is not having PagerDuty configured. Tools are part of the landscape. They are not, by themselves, readiness. I've worked with organisations that had a full enterprise monitoring stack and no functioning incident response process. The alerts fired reliably. Everything else was improvised.
Incident-ready is not the same as having a policy document. A policy that describes what should happen during an incident is not the same as a team that knows what to do during an incident. The gap between the document and the behaviour is where most incidents expose themselves.
Incident-ready is not having had a major incident and surviving it. Survival is not evidence of readiness. Many organisations survive major incidents despite their process, not because of it. Heroic individual effort is not a scalable incident response capability.
What Incident Ready Actually Is
A 50-person company that is genuinely incident-ready has the following in place and functioning - not just documented.
A shared severity framework that the whole team uses consistently.
Every person who might be on call knows the definitions of P1, P2, and P3 well enough that classification is instinctive. When something breaks at 2 am, the on-call engineer can determine the severity and act accordingly without consulting anyone or second-guessing the threshold.
Named roles that are filled without discussion when an incident is declared.
Incident commander. Technical lead. Communications lead. Every person on the on-call rotation knows which role they fill, who their backup is, and what each role is responsible for. When a P1 is declared, within two minutes, there are three named people filling three named roles. No discussion. No negotiation.
Communication templates that are accessible, up to date, and used.
Five templates, stored somewhere reachable in thirty seconds on a mobile device without a VPN. Reviewed quarterly. Used in every significant incident - not just the ones where someone remembers to reach for them.
The test for this is to ask any member of the on-call rotation where the communication templates are and watch how long it takes them to find them.
An escalation path that is specific, up to date, and tested.
Named individuals at each escalation level. Direct mobile numbers. A path that has been tested - meaning someone has actually called those numbers in a non-emergency context to confirm they're correct and that the person on the other end knows they're on the escalation path.
An escalation path with outdated contact details is not an escalation path. It is a false sense of security.
A post-mortem process that closes the loop.
Within 48 hours of every P1, a post-mortem happens. Actions are specific, owned, and time-bound. A follow-up checkpoint is in the calendar before the post-mortem meeting ends.
The test: look at the last three post-mortems. How many of the actions were completed by their agreed deadlines? If the answer is less than 80%, the post-mortem process is not functioning.
A tabletop exercise in the last quarter.
Not twelve months ago. In the last quarter. The friction identified in the exercise has been addressed. The next exercise is already in the calendar.
A team that can run a P1 without any single individual.
This is the most demanding criterion and the most important one. If the organisation cannot run a major incident end-to-end without one specific person being available, it is not incident-ready. It is person-dependent.
The test: name the person your incident response most depends on. If that person was unreachable right now, could the team handle a P1 confidently? If the honest answer is no, that dependency is the most significant gap in your readiness.
The Honest Self-Assessment
Use that list as a self-assessment. Not against an idealised standard - against the specific reality of how your organisation would respond to a P1 declared right now, at this moment.
Not in theory. Not based on what the documentation says. Based on what would actually happen.
Where the gap between the documented process and the actual behaviour is largest is where the readiness work needs to focus. That gap is almost always in the human and communication layer, not in the technical layer. The tools are usually in better shape than the process. The process is usually in better shape than the embedded behaviour.
Incident readiness is when the behaviour, under real conditions, matches the process. That is the standard. It is achievable. And the path to it is specific enough to be planned.
Frequently Asked Questions
What are the key indicators of IT incident readiness?
The six key indicators are: a shared severity framework that's applied consistently without discussion; named roles filled automatically at incident declaration; communication templates that are accessible, up to date, and demonstrably used; a tested escalation path with current direct contacts; a post-mortem process with over 80% action completion rate; and a team that can run a P1 without any single individual. Each of these can be tested against real behaviour, not just documentation.
How do I assess my company's incident response maturity?
Run a realistic self-assessment against specific, behavioural criteria - not against what the documentation says should happen, but against what would actually happen if a P1 were declared right now. Ask: Do my team know the severity levels without looking them up? Can they reach the escalation path in 30 seconds on mobile? When did we last run a tabletop exercise? What percentage of post-mortem actions from the last three incidents were completed?
What is the difference between having an incident process and being incident-ready?
Having an incident process means the documentation exists. Being incident-ready means the behaviour under real conditions matches the documentation: the process has been practised enough to be a habit, roles are distributed broadly, templates are accessible and up to date, and no single person is a dependency. Most organisations have a process. Far fewer are genuinely incident-ready.
How long does it take to become incident-ready?
For a company starting from a reasonable baseline, six to twelve weeks of focused effort - building the process, assigning roles, creating templates, and running an initial tabletop exercise - is sufficient to reach a functioning level of readiness. Maintaining and improving that readiness is an ongoing practice, not a one-time project.
If you would like to test your business incident readiness, you can try our link below:
https://bit.ly/incident-readiness-check
