.st0{fill:#FFFFFF;}

Your Minimum Viable Business Has a Maximum Load. Do You Know It? 

 September 20, 2026

By  Jane Frankland

Over the years I’ve sat through a lot of crisis simulations. Different sectors, different scenarios, different rooms, mostly run by people who know exactly what they’re doing which is the part that matters for what follows.

I’ve lost count of how many times I’ve heard someone say, “That’s fine, we’d move to the manual process.”

Everyone nods, because everyone knows there’s a manual process. And often, the scenario moves on. But there’s a question I don’t hear nearly enough, and it’s this:

How much can you actually do manually in the first five minutes? Then in the first hour? Then on day three, when the backlog has built and the people doing it haven’t slept properly since Tuesday? Then in week two?

Not whether the fallback exists, but what it can carry, and for how long. And I didn’t ask that question either, for a lot longer than I’d like to admit.

It isn’t a gap in anyone’s professionalism. It’s a gap in what we’ve learned to test.

What Good Practice Already Asks

Before a crisis, work out your Minimum Viable Business. Decide what absolutely has to keep running, what comes back first, and what you’ll fall back on when the usual path isn’t available.

This is good advice, and it’s also harder to do than it sounds. Ask anyone who has done it and they’ll tell you the difficulty isn’t the analysis. Rather, it’s getting a room full of senior people to agree that their function isn’t in the first wave.

Business continuity practitioners have spent decades building the discipline to make that conversation possible. The vocabulary reflects that rigour: the Minimum Business Continuity Objective in ISO 22301, alongside recovery time and recovery point objectives and maximum tolerable periods of disruption.

And this isn’t peculiar to one standard or one country. NCSC incident management guidance in the UK, NIST in the US (NIST SP 800-3) and NIS2 in Europe all require organisations, in different ways, to think seriously about incident response, continuity, recovery and crisis management.

And in a mature exercise the questions go well beyond the document. What’s stopped? What must continue? Can we run it manually? What information does that need, and where is it if the systems holding it are down? Who has the authority to switch? Those are the right questions, and plenty of organisations ask them.

So this isn’t a piece about organisations that haven’t done the work. It’s about what the work doesn’t yet ask.

The Question That Comes After The Plan

Have you calculated what happens when everything you’ve lost transfers its load onto the minimum you’ve decided to keep?

Your Minimum Viable Business defines what you need. It doesn’t tell you whether that minimum can carry the organisation through the crisis.

And it’s been bothering me since I wrote about Synnovis.

When the ransomware attack on the pathology provider took out normal blood matching across south-east London, clinicians fell back on O-negative, the universal donor type, the one you reach for when you can’t be certain. I was interested at the time in the connection that failed – the join between a clinician’s question and an answer they could act on.

But there was another question from that story, and I don’t think I went far enough with it. O-negative wasn’t a broken fallback. It was an alternative path, and it worked exactly as designed. The problem was that a whole region sent its load down that path at the same time. Use of O-negative in the affected hospitals rose by more than 90%, and by that July national stocks had fallen far enough for NHS Blood and Transplant to declare an amber alert, with the attack named as one of the drivers.

The lesson is clear. Having a fallback isn’t enough. You need to know what happens when everyone falls back at once.

We Shrink The Organisation And Grow The Demand

There’s a paradox at the centre of every recovery strategy, so familiar it’s stopped being visible.

During a major cyber incident we deliberately collapse a complex organisation onto a smaller one. Fewer systems. Fewer processes. Fewer people, because some are unreachable and others are consumed by the incident itself. Fewer suppliers, because we’ve prioritised the critical ones. Fewer communication channels, sometimes because the ones we’d normally use are the very thing we’ve had to shut down.

Every one of those reductions is defensible alone but together they mean we cut capacity at the precise moment when decision volume, backlog, customer contact, regulatory obligation and time pressure all rise.

Considering structural engineering again, their engineers have a way of thinking about exactly this problem, and they arrived at it after a building fell down.

A Tower In Newham

At about 5.45am on the 16 May 1968, a gas explosion in a flat on the eighteenth floor of Ronan Point, a twenty-two-storey tower in Newham, east London, blew out a load-bearing wall panel at the corner of the building.

The floors above lost their support and fell, and the weight of them landing drove the floors below down in turn, until one corner of the tower had peeled away from top to bottom. Four people were killed and seventeen injured. The building was two months old.

Two things about it matter here.

The explosion was small. The inquiry described it as “within normal limits, in the sense that explosions of this magnitude must be expected from time to time in domestic buildings in which gas is used as a fuel.”

What it concluded about the consequence is what matters: “the extent of the collapse subsequent to the explosion was inherent in the design of the building.”

The building complied with the standards of the day. The weakness was compliant.

And there was no second path for the load. Each precast panel carried the one above it, so remove one and everything above was, by design, holding onto nothing.

What interests me is what the profession did next. It didn’t specify stronger panels. It changed the question the design had to answer. Within two years the Building Regulations carried a requirement that hadn’t existed before, and still stands today as Requirement A3: a building must not suffer collapse to an extent disproportionate to the cause. One way of demonstrating it is to notionally remove a supporting wall or column and show that the resulting damage remains within acceptable limits.

Strip away the engineering and you have a discipline that demands proof that when something is lost, the load has somewhere else to go.

We don’t require that proof. We require a plan.

Six Sentences Worth Finishing

Recovery plans are full of sentences that stop a little too early. Here are six I often hear. The first is where this blog started.

  • We’ll move to manual processing… but how many transactions per hour, and does that rate still hold on day three, when accuracy and staffing start to give out?
  • We’ll use emergency communications… but can that channel carry everyone at once, and does everyone actually have it, installed and working, on a device they can reach?
  • We’ll restore Tier 1 systems first… but which Tier 2 and Tier 3 dependencies do those systems quietly need in order to work?
  • We’ll run with a reduced incident team… but how many simultaneous decisions can that group make before it becomes the bottleneck rather than the response?
  • We’ll invoke our critical supplier… but what happens when fifty of their customers invoke them the same morning?
  • We’ll operate our Minimum Viable Business… but has it been tested under everything arriving at once, or against one scenario at a time?

Let’s look at the critical supplier one, because it’s one we don’t test nearly enough. A contract may tell you what support you’re entitled to but it doesn’t necessarily tell you what capacity will be available when many customers need that support in the same morning. That’s the O-negative problem in commercial form, present in a signed agreement somewhere in your organisation.

The Load Arrives Before Recovery Does

And the timing matters, because the load doesn’t wait for recovery to begin.

The NCSC’s guidance on recovering from a highly disruptive cyber attack separates the first few hours – contain, assess the damage, establish command, work out your real operational state – from the recovery programme that runs over the following days and weeks, and the rebuild after that. It’s blunt about the sequencing: “Rushing recovery before understanding the incident can significantly increase risk.”

Which means the heaviest demand can arrive before anyone has the situational awareness to direct it confidently. Every fallback invoked in those hours is invoked by people who don’t yet know how long they’ll need it.

Maximum Crisis Load

I’ve started calling it the Maximum Crisis Load. It isn’t a framework and it doesn’t need an acronym. It’s just a name for the demand that lands on whatever is still working, at the moment you have least capacity to absorb it and least information to direct it.

And “whatever is still working” includes people.

The calculation doesn’t need to be complicated.

Take one fallback and work out its actual capacity. How many transactions can the manual process handle? How many calls can the emergency line take? How many decisions can the reduced incident team realistically make? How many practitioners can your supplier actually give you?

Then work out the demand you’d put on it if everyone who needed that fallback needed it at the same time.

Compare the two.

If you know both, you have an engineering judgement to make. If you only know what you need the fallback to do, you have a hope. And if the demand is greater than the capacity, you haven’t found a flaw in the plan. You’ve found the point at which the plan stops being true.

Now if you’re running cyber crisis exercises for a living, I expect you’re thinking, you can’t model every combination, simultaneity is combinatorial, and the scenarios are effectively unbounded.

I think that’s right. Which is why I’m not asking for a model. I’m asking for one number per fallback.

How many transactions an hour your manual process can sustain doesn’t depend on whether the cause was ransomware, a failed migration or a supplier going dark. The scenario is what happened. The capacity is what the fallback can carry. And if nobody can state it, the gap is real whichever scenario turns up.

Capacity Is Not Existence

A capability that cannot carry the load isn’t a capability you have when you need it. It’s a capability on paper.

The same goes for people. A reduced decision-making team under maximum load is not a smaller version of your leadership. It’s a queue, and queues don’t fail visibly. They just get longer while everyone in them believes they’re keeping up! At least that’s what happens in the UK, where we still to queuing rules.

Which is why I keep coming back to alternative load paths. They are preserved choices. When engineers require an alternative path for the load, what they are really requiring is that the remaining structure still has options after something has gone wrong. Not that a Plan B exists somewhere, but that what’s left can redistribute what Plan A was carrying. That’s the whole of what I mean when I say we need to design the whole rather than assure the parts.

I used to say that resilience is the disciplined preservation of choice. This year, I’ve changed my mind. I think that’s survivability.

Resilience is our ability to withstand disruption, adapt and recover. Survivability is what gives us options while we’re doing it. An alternative path when the normal one has gone. Enough capacity to carry the load when demand moves somewhere else. The authority to make a decision when there isn’t time to escalate it.

Survivability lowers the impact, preserves choice and enables you to recover better.

One Thing To Do ASAP

Don’t redesign the plan. Take one fallback out of it and write down its actual capacity. Transactions per hour. Calls per hour. Decisions per day. Seats. Licences. Practitioners on call.

Then count the business areas that would fall back onto it in the same bad morning, and work out what each of them would send its way.

If nobody can supply that first number, stop there. You don’t need the second. That’s the finding.

Ronan Point didn’t change because panels got stronger. It changed because the profession started demanding proof that the load had somewhere else to go.

Now I Want To Hear From You…

My question isn’t the one you’re bracing for. I’m not asking about the fallback that worries you. You know that one, and you’ve probably already raised it.

I’m asking about the other one. What’s the fallback you see organisations rely on that rarely gets questioned? The one everyone assumes will be there, but almost nobody seems to ask how much it can actually carry?

Head on over to LinkedIn and tell me in the comments. It’s where this conversation is happening and I’ll be waiting for you.

Did you enjoy this blog? Search for more blogs that you want to read!

Jane frankland

 

Jane Frankland MBE is an author, board advisor, and cybersecurity thought leader, working with top brands and governments. A trailblazer in the field, she founded a global hacking firm in the 90s and served as Managing Director at Accenture. Jane's contributions over two decades have been pivotal in launching key security initiatives such as CREST, Cyber Essentials and Women4Cyber. Renowned for her commitment to gender diversity, she authored the bestselling book "IN Security" and has provided $800,000 in scholarships to hundreds of women. Through her company KnewStart, and other initiatives she leads, she is committed to making the world safer, happier, and more prosperous.

Follow me

related posts:

Leave a Reply:

Your email address will not be published. Required fields are marked

{"email":"Email address invalid","url":"Website address invalid","required":"Required field missing"}

Get in touch