Your Jira Doesn’t Need to Survive the Storm. It Needs to Keep Serving Eggs.

The World Cup gave the rest of the planet a few weeks to discover that there’s still some good left in the United States. Most of that discovery happened in stadiums. My favorite parts were the short videos of people traveling between them, discovering how big the US actually is and stumbling onto some of the more interesting stops between point A and point B. I still laugh at having to explain that the hour-long bus I took between Galway and my airport in Ireland wasn’t that bad, because for me that’s a standard commute around Atlanta.

Among those discoveries was Waffle House, which a lot of them were seeing for the first time. Go during the day and you get the tame version. Eggs, hashbrowns, a laminated menu with color photos that hasn’t meaningfully changed in forty years, and a server who calls you hon. Go at 2 AM and you get the entertaining version, which I’m not going to describe here. But if you’ve got any free time, and you don’t already know, look up some videos about it.

But neither of those is the best time to visit a Waffle House. The best time is during a disaster. I’ve unfortunately been through a few of those, including one where a mile-wide EF5 tornado took out the transmission towers feeding my town and left some of us without power for almost two weeks. That’s exactly when you want a Waffle House, because it has business continuity down to a science. When Hurricane Michael left several Panama City locations too damaged to reopen, Waffle House parked its food truck in front of one of them and handed out free meals. That’s why the United States government uses it as a gauge for how bad “bad” actually is. Waffle House will find a way to stay open unless it’s literally impossible.

So let’s take a look at business continuity and what lessons we can learn from the aforementioned breakfast chain.

But It’s Cloud, Right?

The second I suggest planning for Jira to be unavailable, a Cloud admin somewhere closes the tab. Atlassian manages the instance, right? That was the entire pitch! No servers, no patching, no 2 AM upgrade window, no sysadmin required. As much as that last part still irks me, it’s largely true.

And I want to be fair about it, because they have (begrudgingly) earned this one. They’re better at running Jira than most of us were, myself included. Disk pressure, JVM tuning, a failed upgrade at midnight, the database that filled up on Christmas Eve. That entire category of failure is gone, and I’m not nostalgic for a minute of it.

But look closely at what that sentence actually covers. Atlassian took over the instance. They didn’t take over everything that could disrupt people using your instance. And to be clear, most of those disruptions may not be your fault, and may not be your problem to solve. (We’ll go into some of the specific failure modes here in a bit.)

But you’re still the one answering the question of why the company can’t use Jira, which means you’re the one in charge of coordinating whose job it is to fix and seeing that it gets resolved quickly, and if that’s not possible, finding a way to keep business going in the meantime.

The Index, Briefly

In 2004, while helping with recovery efforts after Hurricane Charley, Craig Fugate of Florida’s Division of Emergency Management noticed that a surprising number of Waffle Houses were still open and serving food. That led him and some colleagues to develop a scale of how badly an area had been hit, based on the state of the chain:

  • Green. Full menu. Damage is limited and the lights are on.
  • Yellow. Limited menu. Generator power at best, low supplies.
  • Red. Closed. Severe damage or unsafe conditions.

It wasn’t until Fugate became FEMA Administrator and used the term after the 2011 Joplin Tornado that it really started gaining traction. Worth emphasizing, though, that this isn’t an official FEMA program. There’s no published FEMA methodology. It’s a rule of thumb from an administrator that turned out to be useful enough to stick. Which, to be fair, is how all the best rules of thumb are created.

The Part Everyone Skips

Here’s the part that actually gets skipped, and it isn’t one of the colors. The storm doesn’t hand you your color. Your preparation does.

Lose refrigeration and a breakfast chain is finished, unless somebody staged eggs and ice before landfall. Lose power and the waffle irons are dead, but the gas griddle isn’t, which is why the limited menu is eggs and hashbrowns and not waffles. Lose the network and card payments stop, unless there’s enough cash and change in the drawer to keep taking orders. Every one of those is the difference between a red reading and a yellow one, and every one of them had to be decided long before the sky turned.

The Index only works as a signal because Waffle House did the work in advance. It has a hurricane playbook that specifies what to serve in each failure mode: gas but no electricity, or a generator but no ice. It keeps portable generators staged. It has a mobile command center, an RV named EM-50 after the vehicle in Stripes, which I’m including purely because it delights me.

None of that is a plan to keep everything running. It’s a plan to keep working well enough, which is to say a plan for a minimum viable restaurant. The question isn’t how to restore the whole operation. It’s what it takes to keep the lights on (even when the lights are literally off) and keep cooking. And in a part of the country with admittedly a lot of disasters, when that griddle might be the only hot food for miles, it’s the difference between people eating and people going hungry.

My favorite detail comes from Hurricane Irene, when the Weldon, North Carolina store reopened at dawn without power. The district manager handed staff printed grill-only menus, and servers steered customers toward sausage instead of bacon. Not because sausage is better. Because four sausage patties fit on the griddle in the space of two bacon slices, and griddle space was the binding constraint that morning. They also boiled water on the gas grill and poured it through the coffee machine over beans that had been ground before the power died.

That isn’t improvisation. That’s a dependency graph, worked out ahead of time and printed on a card.

And the honest part of the story, the part that should make this feel achievable rather than intimidating: Waffle House didn’t have any of this until Katrina in 2005. Seven restaurants destroyed, a hundred more shut down, and the ones that reopened fast were swamped with customers. The program came afterward, built because a disaster proved it was necessary. Nobody starts with the playbook.

What Can Possibly Go Wrong?

I can hear you thinking. You’re saying, “So, that’s a great lesson on how to keep a restaurant running in the worst conditions, but what does this have to do with Jira and Confluence?” (I keep telling you I can read your minds…)

I’m going to say this again. Atlassian took over the instance. They didn’t take over everything standing between your users and that instance, and they didn’t take over what happens to it once you’re inside. Here’s the cleanest way I know to put it. Your identity provider has an outage. Atlassian’s status page is a wall of green, every service healthy, your data perfectly intact, your site running exactly as advertised. And not one person at your company can log in. That’s a total outage, Atlassian did nothing wrong, and there’s nothing for them to fix.

But your users are still blocked from accessing their services. No matter what Atlassian’s status page says, to your users this is a very real outage. Which means you’re forced to react – trying to figure out what went wrong, what needs to happen, who to contact, how long it’ll take to resolve, and how to get people back to work as soon as possible. And all the while hearing from everyone – including your management chain – about when people will be able to work again. That’s putting you on the back foot before you even start reacting, which doesn’t lead to the best outcomes.

This is a red event. What we’re doing is looking for reasonable, real-world failure modes. And just like at a Waffle House, we need to figure out what steps to take in preparation to turn this into a yellow event. Here’s what yellow looks like on the same morning.

Your IdP is still down. You can’t fix that. But you have a break-glass admin account that doesn’t authenticate through your IdP, with its credentials stored somewhere that isn’t Confluence, because Confluence is behind the login that just broke. So you can still get into the admin console and confirm with your own eyes that Atlassian is fine.

You have a written check order, so instead of forty-five minutes of guessing you spend five minutes establishing that it’s the IdP and not the network and not Atlassian.

You have somewhere to say so that doesn’t require the thing that’s down. A text tree, a phone bridge, a status page on unrelated infrastructure. Decide which one now, because you can’t announce an authentication outage using a tool that requires authentication.

You have a name and a number for whoever owns the IdP, and you know their escalation path, because you asked for it on a calm Tuesday instead of at 8 a.m. on a bad one.

And you have a fallback for the short list of things that genuinely can’t wait. Incident intake moves to a phone number and a shared mailbox. Somebody with named authority can approve an emergency change without the workflow. Both get written down as they happen, and both get reconciled back into work items once you’re green.

None of that stops the outage. All of it turns “nobody can work and nobody knows why” into “we’re on the limited menu for a few hours.” That’s the whole distance between red and yellow, and every item on the list had to be decided before the storm.

Now notice something about that list. Only the first item has anything to do with identity providers. The check order, the out-of-band comms, the named contact, the fallback for critical work, and the person who gets to declare it are the same five answers for nearly every failure mode below. You build the card once, and most of it carries.

Seven Ways You Actually Lose Jira

As I mentioned both with Waffle House and the above instance, you don’t plan for an event. Instead, you are isolating the individual failure mode. This would be like Waffle House not planning for the tornado; they plan for the power outage caused by the tornado. This allows you to be more flexible at the moment because a power outage can be caused by many things, but the playbook is the same no matter if that fallen pole is caused by a wrecked truck or a hurricane.

So what are some realistic failure modes that you might experience on a Jira Cloud instance. These seven should absolutely take part of your preparation list, though to be clear this list is not exhaustive, and your details will (not may) vary from what I have.

  1. Nobody can authenticate. IdP outage, expired SAML certificate, an authentication policy change that locked out more people than intended, or SCIM deprovisioning that ran a little too enthusiastically.
  2. Nobody can reach it. ISP or network outage, proxy failure, DNS, or an IP allowlist somebody tightened on a Friday afternoon.
  3. Work cannot get in. Service desk email intake stops. Customers believe they filed tickets. Nothing exists.
  4. A capability quietly disappears. A Marketplace vendor has an outage, an app license lapses, or an app update changes a behavior your workflow was standing on.
  5. The instance is healthy but wrong. A permission scheme change, a deleted field, a workflow published with a broken transition, or an automation rule doing something enthusiastic at scale. And there’s still no undo, single or bulk.
  6. You lose the keys. Billing lapses, or the last person holding org admin leaves and nobody noticed the bus factor was one.
  7. Atlassian is actually down. The one everybody pictures.

One of those seven is Atlassian’s to resolve. You’re exposed to the other six exactly as much as you were on Server, and to a couple of them more, because the levers you’d once have reached for aren’t yours anymore. And this isn’t me being contrarian for sport. Atlassian publishes a shared responsibility model that draws the line in roughly the same place, putting users and user accounts, Marketplace apps, and the content itself on the customer’s side. Worth noting that theirs is a security shared responsibility model. As far as I can find there’s no availability equivalent, and I suspect that absence is a good part of why so many of us quietly assume uptime is entirely their problem.

If you want to settle this for yourself in about five minutes, go pull your last three Jira incidents. Not outages. Incidents. Times when people couldn’t do their work in Jira. Then count how many were Atlassian’s fault. For most shops, the answer is zero.

What Atlassian Gives You to Work With

Ask most Jira admins about disaster recovery and you’ll get RTO, RPO, backup cadence, and restore procedure. All good. All about getting back to green. Ask what runs at yellow while DR does its thing and you’ll usually get a pause at best.

Before you write your own answer to that, it’s worth knowing what Atlassian actually gives you to build it with. I went through their current documentation, and the honest answer is less than you would hope.

There’s no degraded mode switch. I checked, because I desperately wanted to be wrong about this. There are two levers, and both are repurposed from other jobs. You can build a permission scheme that grants only Browse Projects and swap it onto a space, which is Atlassian’s own documented method and exactly as manual as it sounds. Or you can archive the space, which keeps work items reachable by direct link but not editable, though that is Premium and Enterprise only, it pulls the items out of search, and anyone who had API access before archiving keeps it. Neither is a degraded mode. Both are an admin hand-building one at the worst possible moment. And to be fair, no work tracker could really ship this, because what belongs on your limited menu is specific to your business in a way no vendor can guess.

The things you’d lean on aren’t in the backup. Automation flows, third-party apps and their data, the Opsgenie-powered parts of Jira Service Management including alerts and on-call schedules, Assets for JSM, app access settings, and Jira Product Discovery views, insights and vote fields are all excluded from the export. Read that list with an incident in mind. Your on-call routing and your automation are on it, which means the machinery you would lean on to run degraded is the same machinery that doesn’t come back with a restore.

Whether you can reconstruct the incident was settled months ago. Organization-level audit logging requires an Atlassian Guard subscription or an Enterprise plan, and upgrading doesn’t recover events from before you upgraded. So the audit trail you’d need to explain what happened only exists if you were already paying for it before it happened. That’s a continuity decision wearing a licensing decision’s clothes.

Build the Continuity Map

Early in my career, I was sitting at lunch when a buddy of mine, a developer lead, sat down. I’d just had a rough morning. To be clear, this was back on a Server instance, and I’d spent most of it dealing with an outage. So he starts in, half asking and half serious, about whether I was absolutely sure that outage couldn’t last just a bit longer so the team could have a long lunch. I didn’t take offense, but we did get to talking about why I can’t just leave Jira down without a good reason.

My logic was this. The average software engineer’s salary was $100,000, in 2015 dollars. Fifty-two weeks at forty hours is 124,800 working minutes, so call it $0.80 a minute. At 2,000 engineers, that’s $1,600 of wasted productivity per minute, which puts an hour of outage somewhere around $96,000. And that’s before I get into everyone else who was using the Jira instance.

And that was in 2015. Today I’d have to factor in things like broken automation, lost access to clients, and a thousand other things.

Here’s the problem with that number, though. It assumes all 2,000 engineers stop working the moment Jira does, and they don’t. Some teams are dead inside ten minutes because incident intake runs through the service desk. Others won’t notice until they go to groom a backlog on Thursday. A flat cost per minute is the easiest figure to produce and the least useful one to plan against, because it tells you the outage is expensive without telling you which part of it to fix first.

That’s what the Continuity Map is for. It sorts your work by how long it can be gone before somebody notices, so you know which minutes are the expensive ones.

Take your workflows and sort them by how fast their absence hurts. Not by how much people like them. By elapsed time until something breaks that somebody outside your team will notice.

  • Two hours down. What is already broken?
  • Eight hours down. What is broken now?
  • Forty-eight hours down. What is broken now?

For most organizations the two-hour column is short and specific. Incident intake for a software or ops team. Change approval if you are in a regulated environment. On-call routing and alerting integrations. Customer-facing service desk triage. That column is your eggs and hashbrowns.

Eight hours is where coordination starts to hurt. Release trains, cross-team dependencies, and anything a stakeholder expects a status on before the end of the day. None of it is on fire at hour two. All of it is a problem by hour eight.

Forty-eight hours is planning and evidence. Roadmap work, retrospectives, the metrics somebody reports upward, and the audit trail you suddenly can’t produce when someone asks for it. That last one feels safe to ignore right up until the person asking is an auditor.

Everything else is the full menu, and you’ll get back to it.

The trick, and this is the Waffle House lesson rather than a Jira one, is that you’re not picking the most popular items. You’re picking the ones whose dependency chain survives. The griddle runs on gas and doesn’t need refrigeration, which is why eggs stay on the menu and milkshakes don’t. So walk each capability backward: what dies if a specific app is unavailable? If SSO is down? If automation is disabled? If your Assets data is unreachable?

Most orgs have never drawn that graph. The 2/8/48 framing is useful precisely because it forces the graph without requiring anyone to sit down and draw a graph.

What you end up with is one page, and here’s the test for whether you actually have one. Can you hand somebody a single sheet that says here is what we can do right now and here is who declares it?

That second half is the one people forget. Who has the authority to call degraded mode without waiting on a CAB meeting that can’t convene because the tool that schedules it is down? If the answer is “we’d figure it out,” you’ll be figuring it out at 2 AM with an audience.

To save you the blank page, I’ve put together a printable template with seven worked examples, one per failure mode, so you can see what a filled-in card actually looks like before you write yours. Two caveats before you do. You can’t fill it in alone, which I’ll get to shortly. And it goes stale, so put a review date on it, because a continuity plan describing an org chart from two years ago is worse than not having one at all. You’ll trust it.

And never underestimate actual three-ring binders filled with paper. They’re available when you have no network and no power. Better yet, two of them, in different buildings. Then it’s available even if you suddenly find yourself with one less building. Waffle House prints its own in advance. That’s the entire trick.

Your Limited Menu Probably Is Not Jira

If anyone is going to upset anyone during this post, it will be this section. Especially if you identify with a giant “A.” But here we go. When you’re against the wall and you need to keep things running, you might be forced to get creative. Let’s take Waffle House – they don’t stop serving when the fryer dies. They serve what the griddle makes. Well, unfortunately if Jira is the fryer, I hate to say you might have to select a griddle from another buider.

That is to say if Jira is down, your options may have to be a shared mailbox, or post-it notes, or (dare I say it) Excel. I know, perish the thought, but the whole point of this exercise is to think creatively about what your options are to keep things working. And if the next best tool after Jira is a spreadsheet, you may need to use a spreadsheet – at least until you get Jira back up and things are green again. The menu is the fallback plus the path home. It isn’t a subset of Jira features.

Which surfaces an uncomfortable second-order problem. A single source of truth is a single point of failure with better marketing. A lot of organizations spent the last decade deliberately dismantling their non-Jira process knowledge in the name of consolidation, and that dismantled knowledge is exactly what would have been their griddle. If nobody remembers how change approval worked before it was a workflow, you don’t have a fallback. You have an outage.

Who Actually Owns This Document

If you take one thing from this post, take this one. The person responsible for managing recovery and the person responsibly for business continuity might not be the same person. Recovery is an admin problem. Backups, restores, tooling, sequencing. That is our job and ideally we are good at that. However, recovery is a fundamental different set of problems that business continuity. Continuity is a process owner problem, and most of the answers aren’t in Jira at all. Which approvals can be deferred versus which are regulatory. Who is allowed to accept risk on a change while the system is down. What the service desk tells customers. Whether “we’ll log it later” is acceptable for your compliance posture or a finding waiting to happen. We are stakeholders – and we definitely need to be involved in the answer.

Waffle House’s limited menu was written by operations, not by IT. If you’re sitting down to write your Jira Continuity Map alone as the admin, you’re writing the wrong document, and you’ll discover that during the incident rather than before it. Your job is to bring the constraints. Their job is to make the calls.

Wrapping Up

Most Jira disaster recovery planning targets a return to green. That work is necessary and you should keep doing it. But it leaves the yellow state undefined, and yellow is where you’ll actually spend the incident.

Sort your workflows into 2, 8, and 48 hours. Trace each one back to what it depends on rather than to how much people like it. Write down the fallback and the path home. Name who declares it. Get the process owners in the room, because half those answers aren’t yours to give. Then print it, before you need it.

The admin who has that document already written is the one their organization calls first. Not because they can restore faster than anyone else, but because they’re the only person in the building who can answer “what can we still do?” without guessing.

I’d genuinely like to hear what lands in your two-hour column, because I suspect it varies more by industry than any of us assume. If you’ve been through a real Jira outage, I’m even more curious what you thought was critical beforehand and turned out not to be.

Until then, this is Rodney, asking: have you updated your Jira issues work items today?


Enjoyed this one? These posts are free and always will be — but if this saved you a headache or taught you something worth keeping, you can drop a tip in the Ko-fi jar. No paywall, no subscription, no catch. Just a thanks if the writing earned it.

→ Support The Jira Guy on Ko-fi


Discover more from The Jira Guy

Subscribe to get the latest posts sent to your email.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.