Skip to content

Life. Adventure. Consulting. Technology.

Technology

The Systems Are Down. So What Can You Still Do?

A car maker's plants have now been stopped for a fortnight with its systems deliberately switched off. Continuity planning answers how quickly systems come back. The harder question is what an organisation can still do while they are gone.

Chris Cooper 4 min read
An empty conveyor line running through a large industrial hall in bright daylight

Jaguar Land Rover’s plants have been stopped for a fortnight. The systems were taken down deliberately to contain an attack, and the restart date has already moved more than once.

Most of the attention is on how it happened. The question worth taking from it is a different one — what an organisation can do while the systems are off.

1. Continuity Answers How Fast, Not What Instead

Continuity planning is built around recovery time. How long until the service is back, how much data is lost, what is restored first.

For most incidents that is right. A failed server, a bad release, a lost data centre — they come back, and the only question is the interval.

It stops being right when they are off for weeks, or running with nobody willing to trust them. There is no interval to plan for, only what can be done instead.

2. The Manual Fallback Was Part of the Saving

Every system that replaced a manual process removed it. That was not an oversight — it was the business case.

So the paper form went, and the local register, the printed list in the van, the person who knew the round by heart. All correctly retired.

What went with them was the ability to work without the system at all. Nobody counted that, because nothing came to collect it for twenty years.

3. What an Outage Investigation Actually Finds

I was once part of a team investigating a major outage at an airline group, where the failure spread to its other carriers and grounded flights for days.

It turned out to be a power problem. Nobody knew that at the time — and every recovery estimate given over those days rested on a cause nobody had established.

Meanwhile the disruption was made downstream: where passengers were, which crew could legally fly, which bags were where. The aircraft were serviceable, the staff were at work, and nobody could tell them what to do.

4. The Cost Lands Outside Your Building

An extended outage does not stay inside the organisation that had it. Suppliers stop being paid and stop being told what to build — and a good number cannot carry a month of that.

Those firms have done nothing wrong and no say in the recovery. Their exposure is set by how long somebody else takes.

A continuity plan that stops at your own perimeter is measuring the wrong thing. The real number is how long the businesses around you can survive your outage.

5. Emergency Services Rehearse the Bad Version

The organisations that handle this best have no option to stop.

I have taken part in several paper operations exercises in the NHS — the digital systems deliberately stood down, the service running on continuity processes instead. As an emergency service, the work does not pause while somebody restores a database.

They are uncomfortable, and that is their value — the only reliable way to find out which parts of a plan describe something people can do.

6. A Document Is Not a Rehearsal

Most organisations have a continuity plan. Very few have run one. It was written for an assessor, and it reads well.

Run it and the gaps arrive within the hour. The form names a role that no longer exists. The signatory list is four years old. The contact details live in the system that is down.

None of that is discoverable on paper. Only by making people do it on an ordinary afternoon with the systems off.

7. What Is Worth Establishing First

Not everything needs a manual alternative, and building one for everything would cost more than the outage. The narrower question is what is needed on the first morning:

  • which processes have a version that works without the system, and when it was last run
  • who may decide things while the approval workflow is unavailable
  • where the information people need exists outside the system that holds it
  • which suppliers and which customers cannot absorb a month of this

The last is usually the one nobody has looked at — and the one that turns an internal incident into somebody else’s insolvency.

8. Say Which Kind of Carrying On You Mean

There are two versions of carrying on. Organisations plan for one and experience the other.

The first is the systems coming back. That has an owner, a budget and a tested procedure, and most manage it well.

The second is the service continuing while they are gone. That has none of those things, and no date where anybody finds out.

Final Thought

Recovery is a technology question, and it is being answered. Somebody owns it, funds it and tests it.

What the organisation can do meanwhile is a management question, and in most places it has never been put. It gets answered anyway — on the morning it matters, by whoever is standing there.

Stay in touch

Occasional writing, straight to your inbox

A short note when I publish something worth reading. No noise, and easy to leave whenever you like.