Managing an incident

Power’s out
On a recent trip we stopped to eat at a chain restaurant, one of those with a drive-through, self-service ordering stations and such. The usual was taking place: somebody selected the wrong meal and had to start over, somebody didn’t know what to order and made everybody wait while they looked through every single item on the menu.
As we were assembling the order the screen went black. And the lights above us. And the music stopped. A quick glance towards the kitchen and I noticed a couple confused and scared faces. The manager shouted: “we had a power cut!”, and quickly disappeared from the view.
10 seconds later the lights on the machines in the kitchen, and those above us, switched back on and we heard “it’s coming back on!”.
The reaction from the team was so good that it deserves to be shared.
Response
As soon as the machines came back on the team gathered around the order assembly table and started drafting a plan. Not everything was fine. The POS kiosks were still off. The kitchen system was dead.
The manager reassured people that the kiosks will get back up in a bit, people simply waited in front of them knowing that it would take a while. It took about 6-7 minutes.
While we waited I could see the team get into a flow. As soon as the equipment was fully operational they completed the orders they had printed before the power went out and completed their preparation. One of the staff approached the customer who had already ordered, explained that the kitchen had no visibility into the orders and asked for the receipts to know what was ordered. These were the next step.
In the meantime the drive-through was shut down. We only learned about it when leaving but it made total sense - there was no way to manage the flow of serving the meals.
A member of the drive-through staff was circling the restaurant apologising for the inconvenience and ensuring that everybody got what they needed.
A customer entered the restaurant with an order made through an app. One of the kitchen staff looked at it and quickly managed the expectations: we had a power cut; the kitchen systems are down; we have no way of seeing your order and I don’t know how to solve this; give me a sec, I’ll figure out what we can do.
When the screens came back up in the kiosks we completed our order. The kitchen screens were still down. I handed my receipt over.
We received our food, one sandwich was missing. I approached the counter, showed which one it was and after a minute received an apology and the missing sandwich.
What I think happened but I didn’t see myself
The kitchen system was down way longer than I would think was usual. I suspect the manager was on a call with internal support, trying to figure it out.
I also suspect they used internal support to get the online orders.
We ate, and I learned
The meal was everything one would expect from such a place. It took longer to get than expected, but it came with lessons on resilience:
- The person in charge focussed on recovering the core operations (making sure the kitchen equipment is up and functional)
- The team came together and planned their response
- They used the backup artifacts (printed orders) and audit trail (receipts) of the system to deliver their services
- They executed a graceful degradation (closed the drive-through, blocked the entrance with cones, repurposed the drive-through staff) of the system to preserve the ability to provide any services
- Customers were informed about the status of the incident and the invisible constraints (kitchen screens down)
- The expectations were managed (kiosks recovery, in-app order, drive-through physically closed)
- Capacity was repurposed to make most impact (a member of drive-through staff checking in on the customers)
Things that work if you prepare
I don’t know if this was something the team trained for or experienced in the past, but they handled it well. They made the calls on what to let go of and what to double down on.
There was uncertainty about how to handle some of the problems, but no doubt about the intent. The intent I read was: to deliver the best service possible given the circumstances and not to leave any customer in the dark (literally). Customers were treated respectfully, the communication was clear and kind. Loss of speed was compensated by proactive care.
There are also the systems that just worked as designed - as soon as the power went out emergency lights came up in the toilets. The recovery was a bit slower - the emergency turned off immediately and the normal lights only switched on after a couple seconds. Not as smooth but also not a big problem. No human involvement needed.
Back on the road
As pleasant as it was to watch the team manage the incident we finished our meals and continued the journey.
It’s good to be reminded that anything can fail, and how we handle it can make or ruin someone’s day, week or life. Prevention is only one of the strategies that can be applied, and sometimes it can be out of your hands or not worth it. Resilience has more options:
- plan for failure,
- identify the critical path and use graceful degradation for the rest,
- organise incident management drills
- bake self-recovery into the subsystems
- learn through the experience
- make the degradation (and the recovery) visible
I would like to recommend a few reads if you want to learn more about the whole spectrum of resilience:
- “Why We Still Suck at Resilience” [2] - helps get better at resilience, and shows the risks once you succeed
- “Observability Engineering” [3] - a guide through growing systems with observability in the centre; can be downloaded from the Honeycomb page or purchased on paper
- “Resilience Engineering” [4] - great article and resource; at the bottom you can click “Explore in network” to find related articles
References
- rawpixel. Desktop, Old, Champagne image. Pixabay. https://pixabay.com/photos/desktop-old-champagne-cloth-3147855/
- Adrian Hornsby. Why We Still Suck At Resilience. Organizational Dynamics. https://leanpub.com/whywestillsuckatresilience
- Charity Majors, Liz Fong-Jones, George Miranda, Austin Parker. Observability Engineering. Achieving Production Excellence. 2nd Edition. https://www.honeycomb.io/observability-engineering-oreilly-book
- Tom Geraghty. Resilience Engineering. https://psychsafety.com/psychological-safety-resilience-engineering/
