A service is restored, the immediate pressure eases, and everyone returns to other work. Then a similar interruption happens again. For a growing business, the cost includes delayed customer work, time spent chasing updates and staff who cannot confidently plan their day.

Recurring incidents deserve a closer look at what happens around recovery. A successful restart or workaround may restore the service while leaving important conditions unchanged.

Start by recognising what recovery achieved

When customers cannot place orders or staff cannot access information, getting the service working is a legitimate priority. The people responding may be making difficult decisions with incomplete information.

Once the immediate interruption is under control, a different question becomes useful: what still needs attention now that the service is available again?

That question creates room for follow-up without treating the response team as the problem. Recovery and longer-term improvement require different evidence, time and decisions.

Similar symptoms do not prove a shared cause

Two interruptions can produce the same error message for different reasons. Equally, apparently unrelated symptoms can involve the same dependency.

Compare what was happening around each event. Was a supplier sending information late? Had a change occurred? Was the workload different? Did the issue appear during a handover or when a particular person was away?

Keep observations separate from explanations. “The queue stopped after the update” is an observation. “The update caused the failure” needs further evidence. Avoid committing to a technical fix before the team understands which explanation it is testing.

Ask what remained after the fix

A few questions can reveal work that recovery did not resolve:

A task marked “investigate” can remain open indefinitely if nobody knows what decision it needs to support. Make the next action specific enough for someone to complete or escalate.

A hypothetical example: restarting an order integration

This is an illustrative scenario, not a Neith client case.

An integration transfers orders from an online shop into the system used for fulfilment. It stops processing, and a restart gets the queue moving again. The technical alert clears.

Customer service still has unanswered questions. Were any orders missed? Did a retry create duplicates? Who checks that the fulfilment system now matches the orders customers placed?

The restart may have worked exactly as intended. It does not, by itself, establish that the business service has fully recovered or explain why processing stopped.

A useful follow-up separates three tasks: investigate the interruption, reconcile affected orders and decide what evidence will show that the corrective action helped. Those tasks may belong to different people. Someone still needs to coordinate the overall follow-through.

Compare a small set of incidents

Start with a manageable sample of recent events affecting one service. Use existing records and conversations; a new reporting system is not a prerequisite.

For each event, capture:

  1. The service and business effect. What work stopped or slowed, and who was affected?
  2. The observed symptom. What did people or monitoring actually detect?
  3. The operating conditions. What was different, and which dependencies were involved?
  4. The recovery action. What restored operation, including any manual steps?
  5. The remaining action and owner. What still needs doing, by whom, and what is blocking it?
  6. The verification. How will someone check that the action changed the relevant condition?

This is a starting point for discussion. Where records are incomplete, mark the gap rather than filling it with a confident explanation.

Check whether the action changed the situation

Closing an action and establishing that it helped are different things. Agree what evidence would support the expected effect before considering the work complete.

For example, if the concern involved delayed supplier data, check arrival times and how late files are handled. If manual reconciliation was the gap, check whether someone can identify and resolve incomplete transactions using the agreed process.

Choose an observation period that includes the conditions associated with the problem. A few quiet days do not demonstrate that a month-end issue has been resolved. The absence of another incident is useful context, but it may not be sufficient evidence on its own.

When the wider service needs attention

A broader review can be useful when actions repeatedly stop at team boundaries, several suppliers are involved, or recovery depends on the same few people. These are operational questions as well as technical ones.

Clear end-to-end ownership helps people decide who coordinates the concern, who can authorise action and how unresolved issues are escalated. Neith’s approach to recurring incidents looks at those connections around the service.

Take these questions to the next review

Start with one recurring problem and make the follow-through visible. If the concern crosses several responsibilities and remains difficult to explain, the Service Reliability Health Check offers a way to explore it within an agreed scope.