
Why Now Is the Best Time to Improve Your IT Incident Process
Why the Best Time to Fix Your Incident Process Is When Everything Is Fine
Throughout my career, I've watched a pattern repeat itself across organisations of every size and industry.
Something goes wrong. A major incident. A near-miss. A client relationship that was damaged by a poorly handled outage. In the aftermath, urgency sets in. The board wants action. The CTO wants to fix it. Budget is found. Time is allocated. Work begins.
And then, gradually, other priorities return. The sprint fills back up. The urgency of the incident fades as the memory of it fades. The process improvement work gets 80% done and then stalls.
Three months later, the organisation is in approximately the same position it was before the incident.
Until the next incident makes the conversation urgent again.
Why Urgency Is the Worst Condition for This Work
The instinct to fix things after they break is deeply human and entirely understandable. It is also one of the least effective conditions for process improvement work.
When urgency is driving the work, the scope tends to narrow to the specific failure that just occurred. The runbook that didn't exist for this specific failure mode gets written. The escalation path that wasn't followed gets reviewed. The missing communication template gets created.
What doesn't happen is the broader structural work. The comprehensive review of all the runbooks, not just the one that was needed. The establishment of a regular tabletop exercise programme. The role clarity across the full incident severity spectrum. The cultural work around escalation and blame, which doesn't feel urgent after a specific technical failure.
The urgency produces a targeted fix for the specific problem. It rarely produces the systemic improvement that changes how the next, different incident is handled.
What Calm Conditions Make Possible
When everything is fine - when the last incident is far enough in the past that the urgency has faded, when the team has capacity to think rather than react - the conditions for process improvement are completely different.
There is time to do the work properly. The runbooks can be written comprehensively rather than reactively. The tabletop exercise can be designed to test multiple failure modes rather than just the one that most recently occurred.
There is the right emotional environment for honest assessment. The conversation about where the gaps are does not carry the weight of a recent failure. The team can discuss what would go wrong without anyone feeling defensive about what did go wrong.
There is space to practise. The tabletop exercise that identifies friction can be followed by another that tests whether the friction was resolved. The communications lead can practise writing an update in a non-emergency context, where getting it slightly wrong is a learning rather than a cost.
There is the opportunity to build habits before they're needed. A process practised repeatedly in calm conditions becomes a habit when the pressure arrives. A process that was built in urgency and never practised is available only as a document - which, under pressure, is almost the same as not available at all.
The Continual Improvement Principle
The ITIL framework has a concept that is directly relevant: continual service improvement. The idea is not that processes get fixed after they fail - it is that processes are regularly reviewed and improved as a standard practice, regardless of whether a failure has occurred.
This means that incident management is not a project with a completion date. It is an ongoing discipline. The playbook is not written once - it is reviewed quarterly. The tabletop exercise is not a one-off - it is a regular programme. The post-mortem process is applied consistently enough that it becomes part of how the team thinks about every incident, regardless of severity.
Organisations that operate this way have a fundamentally different relationship with incidents. Not because they have fewer incidents. Because when incidents happen, the response is structured and confident rather than improvised and anxious. The damage is lower. The recovery is faster. The client relationship is more resilient.
That difference is not built in urgency. It is built in the quiet periods, through deliberate and sustained investment in process when there is no immediate pressure to do so.
The Honest Calculus
The cost of building and maintaining a functioning incident process in calm conditions - the time invested in documentation, tabletop exercises, role clarity, and regular review - is predictable, bounded, and spread across time. For a 50-person company, it might represent two to four hours per quarter of leadership time, plus a half-day exercise once a quarter for the relevant team. Significant, but manageable.
The cost of building it reactively - after a major incident that damaged a client relationship, triggered an SLA conversation, or required a board explanation - is unpredictable, concentrated, and accompanied by all the costs of the incident itself. It is always higher. Often significantly higher. And it is paid while simultaneously managing the aftermath of the incident that triggered the urgency.
The choice between those two cost profiles is available right now, in this moment, when everything is fine. Not as a moral argument - as a practical one. The proactive investment is less expensive, more effective, and available when the organisation has the capacity to do it properly.
The next incident will happen. The question is whether it finds a process built carefully in calm conditions, or one assembled in a hurry after the last one.
Frequently Asked Questions
Why is it better to improve your incident process before an incident happens?
Calm conditions allow for comprehensive, systematic improvement rather than reactive point fixes. The team has the emotional space for honest assessment, the time to practise what gets built, and the opportunity to embed habits before they're needed under pressure. Reactive improvement produces narrow fixes for the specific failure that triggered the urgency; proactive improvement produces structural resilience.
What does it cost to maintain a good IT incident process?
For a company of 15 to 200 people, the ongoing maintenance investment is modest: approximately two to four hours of leadership time per quarter for review and updates, plus a half-day tabletop exercise once a quarter for the relevant team. The post-mortem process adds roughly 90 minutes per significant incident. This investment is a fraction of the cost of a single major incident handled without a functioning process.
How do I justify investing in incident process when nothing has gone wrong recently?
The absence of a recent major incident is not evidence of low risk - it is evidence of luck holding. The honest question is: if a major incident occurred today, would the response be structured and confident or improvised and anxious? If the answer is the latter, the gap is real and its cost is accumulating quietly. The investment is justified by the cost of the situation it prevents, not by the cost of the last incident.
What is ITIL continual service improvement?
Continual service improvement (CSI) is an ITIL framework concept that treats process improvement as an ongoing practice rather than a reactive response to failure. Applied to incident management, it means regularly reviewing and improving the incident process - updating runbooks, running exercises, refining communication templates - as a standard operational discipline rather than waiting for a major incident to make it urgent.
