What the Hell Happened (and what can I do about it?)
Something happened this week that changed the conditions your organization is operating in. Most leaders will find out too late.
What the Hell Happened? is a weekly 30-minute briefing for executives in healthcare, higher education, manufacturing, logistics, transportation, and construction. Every episode takes the week's most consequential events — policy shifts, system failures, supply chain disruptions, regulatory changes — and works through what they mean for the people making decisions at the top.
Three segments. Every episode.
Readiness — what happened and why it matters. Resilience — what it reveals about the assumptions your organization is running on. Advantage — one specific action before Monday.
Hosted by Mike McCracken, founder of Southwind Planning Solutions, with decades at the intersection of emergency management and private sector operations.
What the Hell Happened (and what can I do about it?)
Episode 2: When Crises Don't Wait Their Turn
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
On August 1, 2012, Knight Capital Group lost more than $400 million in forty-five minutes. Ninety-seven automated error alerts fired before the market opened that morning. All ninety-seven went unread.
Three years later, Target Canada's inventory system reported normal operations while distribution centers overflowed with product that couldn't reach store shelves. Customers walked in, found nothing, and didn't come back. 17,600 employees lost their jobs. The losses exceeded five billion dollars.
Different industries. Different scales. The same failure underneath — two crises arrived simultaneously, and neither organization had a plan that accounted for that possibility.
In this episode, Mike McCracken examines what both failures share, and what the emergency management world has understood about compounding crises for decades. The doctrine exists. Most organizations have never heard of it.
Deep Dive available: bluetogray.beehiiv.com
Southwind Planning website: www.southwindplanning.com
www.southwindplanning.com
www.bluetogray.beehiiv.com
Sometimes it's a policy shift. Sometimes it's a system that fails. Or it could be that a market moves. Often leaders find out too late. The signal was there, but nobody turned it into something they could act on before the moment passed. What the hell happened is a weekly briefing for leaders who would rather be ready. I'm Michael Crackett, and I've spent over 30 years working in situations where planning gets interrupted by reality. This program looks at not just what the hell happened, but also what can be done about it. Welcome back, everybody, to another episode of the podcast. I'm truly excited to have you here with me again, and I appreciate your interest and taking the time to listen to these episodes. It means a lot to me, and I really appreciate any comments that you'd like to leave, and it'll help drive us to future episodes, any topics that you may be interested in hearing about, and also give you a chance to leave a review, let other folks know what you think, and give me a thumbs up if you like what you're hearing, or let me know some things that I can do to improve. It all helps us get better. So anyway, I'm glad that you're here. Let's jump into this episode. The title for this episode is When Crisis Don't Wait for Their Turn. What exactly does that mean? What it means is a lot of plans, most plans in fact, are written around focusing on one crisis at a time. When that crisis turns up, they activate the plan or the plan goes into action and the crisis is dealt with. But what a lot of plans, in fact, most plans don't address is how to deal with a situation that comes up, you're working on it in the plan, and all of a sudden something else goes wrong, and now you're facing two crises at the same time. They may or may not be related and they may or may not be connected, but they will have an impact on each other and definitely have an impact on your organization. So that's what we're going to talk about today. So with each episode, I try to open with a single moment. A blue to gray illusion is what I call it. It's a blue sky moment where the status says that everything was fine in the organization, but in reality the skies were already turning gray. Things were getting dark, and things were happening behind the scenes. So here's this episode's blue to gray illusion. The leadership dashboards showed the inventory was moving normally. The distribution centers were packed with the product, but it couldn't be shelved, and the store shelves were actually empty. The system reported one reality, but the warehouse showed an obvious different reality. The status and all the indicators said that everything was okay, but the reality was much different. So what the hell happened? Well that's what we're going to talk about today. First let's look at readiness. Here's the readiness portion of this episode. Readiness is where we're going to look at what actually happened. We'll talk about the sequence of the events and the decisions that were made, and then we'll look at some of the moment that things went wrong. Today in readiness we're going to look at two organizations and three different incidents actually. They affect different industries, they happened during different decades, and they impacted different people. But they made the same catastrophic mistake in a similar way. So let's start first with a question. You have an organization and they hit a problem. A competent team makes the appropriate response according to the plan. You have confident people and they work fast to try and solve the problem. They run all of their standard corrective procedures and the procedures only make it worse. This doesn't happen because the procedures were slow or the response was delayed, and it didn't happen because the responders or the folks managing the situation didn't care. They were operating on a picture that was based on information that they thought was actually true, when in fact the information that they were operating on wasn't the picture that was actually true in reality. So what happens when the correction is based on the wrong diagnosis? A diagnosis that's based on information that's incorrect, inaccurate, or not timely. Let's look at these two case studies. The first one is Knight Capital. It's an investment group, and the problem happened on October first, 2012. On this date, Knight Capital deployed a new trading software and it went out to eight servers, but one server got missed. It had a legacy code on it called PowerPeg, and that stayed active on that one individual server while the new trading software deployed as it was supposed to to the rest of the servers. At eight oh one AM there were ninety-seven emails that were generated that showed an error. They were automated emails, and they identified the problem. The problem here was, though, that they were not configured as high priority alerts, and so nobody actually read the emails before the market opened. The market opened at 9 30 AM. The unpatched server began to execute millions of erroneous trades automatically. The team seized the problem and they ran their bad deployment response as it was outlined in their plan. It called for them to uninstall the new code from the functioning servers, so that's what they did. And then that action reactivated the legacy code across all eight servers. What happened then was what they thought was the fix actually served to multiply the damage because the old server and the old code was now activated, running on the wrong information. The result was in the time of forty five minutes, over four hundred million dollars in losses were accrued by the company. The deployment wasn't the error that killed this company. That wasn't the critical impact. The critical impact was a corrective action that was built and deployed on a misdiagnosis of what the actual problem was. The actual problem was that the trading software had not deployed on all of the servers, it had missed one. But that was not actually noticed until after what they thought was wrong was addressed in the wrong way. So the team was competent, the team did what they were supposed to, but the picture was wrong. Here's another case. This one involves Target. It's actually Target in Canada. They're called Target Canada. They acquired Zellers, which is another department type store in Canada, and they acquired the leases for one point eight billion dollars. And they committed to opening one hundred and twenty four stores simultaneously in Canada because the real estate costs demanded that they do so. So what happened? They stood up a new ERP system and it got populated manually. This happened under an extreme time pressure due to the deadline to get everything open. At launch, up to seventy percent of the data was found to be corrupted. There were wrong dimensions, wrong barcodes, wrong case pack counts. That happened because it had to be populated manually due to the time restraints and so some of the errors were made by the manual entry of the data. The automated replenishment system didn't know that that data was not accurate, and so it treated it as it was accurate. What happened next was the distribution centers overflowed with the wrong product while the store shelves sat empty. Everything that they needed to order was automated and sent to the distribution centers. However, once it arrived at the distribution centers, it never made it out of there and actually hit the store shelves. So what happened next was the leadership decided that they needed to disable the auto replenishment software as outlined in their plan because something was obviously wrong. And they decided that they would go manual, try to get the supplies out to the stores and make the deliveries. What that resulted in was the manual corrections were now feeding new bad data back into the same corrupted system. So they were still using their distribution system, but it was still relying on the bad data that had come from the improper system that was originally activated. They still had the dashboard that showed one reality, everything was fine. But the people in the warehouse were reporting something entirely different. The warehouses were overstocked, nothing on the shelf in the store. According to the filing on january fifteenth, twenty fifteen, one hundred and thirty three of the stores were closed. This resulted in seventeen thousand six hundred people who were now out of work and over five billion dollars in pre-tax losses for the company. Every corrective action, both automated and manual, was operating on a foundation that was wrong from the start. One other highly publicized incident, I just watched a movie on this as a matter of fact recently, and it really shed a lot of light on it, really a good movie, but it involved Deepwater Horizon and the disaster that resulted from that oil spill. This was a third organization, an entirely different industry, this time the oil and gas industry, not technology and not retail. In this incident, as you recall, the crew of the Deepwater Horizon misread a negative pressure test on the offshore oil rig. They interpreted the results through the lens of their last coherent mental model, and that assumed that the well was fine. The test made sense and they thought that they could proceed without any problem. In reality, the well wasn't sealed. The CSB documented this as a diagnostic failure interacting with a physical safety system failure. The information was actually there, and the information existent. The problem was the process was not followed. But what arrived at each decision point had not been verified against an independent read before the action was taken. You can find out more about this information. There's a lot of public information, a lot of movies and other research, after action reviews, a lot of things on the internet that are available that go into the details, but we're focused mainly here on the processes. In fact, I have a colleague that I work closely with who is a direct responder to this disaster in Deepwater Horizon, worked on the Gulf Coast at the time and was responsible for a portion of the response and the recovery after it happened. I hope to have him on and interview him about this and what the processes that failed as far as the response went. There were other issues that were involved there too, so hopefully I can get him on the podcast on a future episode and talk a little bit more in detail about that. But we'll also look at that more later in resilience. We're going to look at exactly where the gap lives and why it's so hard to see until it's too late. So both of these organizations had smart and capable people, and they were running hard at a real problem. But in the end both had made it worse. The question isn't whether they were competent because they were. The question is actually what information were they actually using to operate on and to base their response? Sit on that for a minute. That's the readiness picture. That's what happened and what it looked like from inside both of the organizations. Now let's move into resilience. Resilience is where we're going to stop looking at the events themselves and start looking at the mechanisms that made the events occur. Because these two failures aren't parallel stories that just happen to resemble each other. They share a similar architecture, and once you see it, you can't really unsee it inside your own organization. In both cases, both Night Capital and Target Canada, the information systems and the operational systems were separated completely before any corrective action was taken. This didn't happen gradually over time, and it wasn't because of negligence that had built up across months. This was a structural separation, and it was present before the first decision in the crisis was even made. And then every decision after that that was made downstream was built or based on a corrupted picture based on bad information. In the case of the Knight team, they had corrected for the wrong failure. Targets leadership had fed new bad data into an already broken system. These were competent people working in a wrong reality. And that's not a technology problem. That's not even a training problem. What that actually shows is an architectural problem, something in the structure that just wasn't working right. I spent thirty years in emergency management, and that teaches you a lot. It teaches you something about how systems can fail under pressure. And it rarely results from steps that were taken in the process that broke. People are always following the process, following the steps, taking the steps that are needed and doing their best. The steps are documented and they're trained and they're exercised. But what breaks is what happens in the space between those steps. This is where the information is moving from one part of the system to the next, and it's critical to making things work as they should. In both of these cases, that's exactly where the failures lived. For Night Capital, the team followed a legitimate, bad deployment response, as they were supposed to and as was outlined in the plan. The steps were correct, but the information that arrived at each step, which was about the servers that were affected, and that's what the root cause was, had not been verified before it was acted on. In Target Canada, the leadership followed the legitimate crisis response for the incident they were dealing with, the problem with the stock shelves in the stores. They identified the problem and disabled the automation system. They went on manual and the steps made sense for what they were dealing with according to what they thought was occurring. The problem was that the information that those steps were built on was corrupted from the start. In emergency management, one of the doctrine standards is real-time, accurate situational awareness. Not just having the information, but maintaining a verified, current picture of what is actually true as the conditions are always changing. The emphasis here is on verified, not reported. Not assumed, but confirmed. I had someone tell me one time there's basically two types of information, confirmed and unconfirmed. If the information's confirmed, then it's safe to act. If it's unconfirmed, then you're assuming a risk. So depending on how you've obtained the information and how sure you are the information is accurate, you're assuming a certain level of risk by taking that as unconfirmed information and making a decision based on that information. Because in a fast-moving situation, the most dangerous thing in the room is a confident action that's built and taken on unverified or unconfirmed information. So before I apply a framework here, I want to name it because it's going to be foundational to everything we're going to talk about. And especially if you're new to this show or you are an occasional listener, I want to give you a 30-second background version to discuss and point out exactly what we're talking about here. What I've come up with I call business lifelines, and these are the critical functions that every organization depends on to keep themselves operating under stress. These aren't theoretical categories, they came out of the same discipline that asks when an organization is under serious pressure, what's going to break first and in what order, and how do we know that? One way to think of this is as the load-bearing walls of your operation. That's really what these business lifelines are. When a lifeline degrades or starts to break down, one of those walls starts to give way and it doesn't stay contained, it still can't support its weight. It puts stress on itself and everything that's connected to it. Business Lifeline gives us a common language for identifying where organizational capacities are and where they're actual failing, as opposed to where the dashboard says that it's failing. If you can monitor the lifelines and monitoring them accurately, and compare that to what your dashboard's saying, it's a much more clear picture of what's actually going on so it can be dealt with properly. In this episode, there are two lifelines that are relevant. Keep that in mind as we go forward. The next thing we need to discuss here is the information hygiene concept, and this becomes a central concept in this incident and why it matters more than most organizations realize. Information hygiene isn't about whether you have data, both of these organizations had data. It's about whether the information moving through your system has been verified, vetted, and is confirmed as accurate before someone takes action. The gap in both of these cases was not at the decision points, it was in the handoffs, or the spaces in between the steps where the information moved from one part of the system into the next, and where nobody checked on whether that information that arrived was actually true and accurate. In the case of Night Capital, if you remember, I said that there were ninety seven emails sent at eight oh one AM. These were automatically generated by the system, so the information existed and the pathway to act on it was taken in the form of an email. However, that pathway was broken, and no one verified that the deployment was complete across all eight servers before the market opened because those emails weren't identified as high priority. That's a pathway failure. The signal was present, but the verification step didn't exist. Neither did the notification that the information was there so the verification could take place. In the case of Target Canada, corrupted data entered the ERP at the foundation at the very beginning. The system processed that information as though it was accurate, so every automated decision downstream, and every manual correction that was attempted by the staff was built and based on information that had never been confirmed against physical reality. That is a foundational failure. The data was there again, but nobody had verified that the data was true and correct before taking action. These are different mechanisms, but they have the same missing step. The missing step is verification at the handoff point before the information becomes a basis for action. So let's look at this from the perspective of how it matters to your own organization. When information hygiene fails, it doesn't produce an information problem that stays in the information system. It actually produces a lifeline problem. In the case of Night Capital, operational execution, which is lifeline number three, was degraded because the corrective actions were based on a misdiagnosis of which server was actually causing the problem. The team was executing correctly against the wrong target. And in Target Canada, Lifeline one, or the supply chain and resource flow lifeline, had collapsed entirely. But the mechanism wasn't a supply chain problem, it was an information system feeding false instructions to the operational system. This happened from the moment that it went live. These two failures aren't parallel. One caused the other, had a direct impact on the other. The information hygiene failure created the conditions for the lifeline failure. This illustrates the direct link and the value and importance of clean information, valid information, because that's what's going to feed the status of the business lifelines to give you a clear picture of what the conditions actually are. Once the information layer was compromised in these cases, every operational decision downstream was built on that bad foundation that had already given way. Something else that I'd like to discuss and that we've mentioned in the previous episode and we'll mention again because it's key is what I call the delta of decay. This is something else that I'd like to discuss here is the delta of decay. In Target Canada, which is an extreme case, the corruption at launch was visible almost immediately, but the mechanism it exposes operates in most organizations on a much slower clock. Data that was accurate when someone entered it, but it hasn't been touched in a while, or processes that worked when they were written, but the team no longer exists, are both things that can linger for a long period of time. This expands that delta of decay. The delta of the Of decay is an indicator that illustrates the time between when something actually went wrong and when something is actually understood as an issue or a problem and addressed. Things like vendor relationships that functioned well until the person who managed them left and is no longer at the company. The gap between what the system reports and what actually is true widens gradually. But nobody flags it because nobody's watching for it. It shows up on the day when someone tries to take action on it under pressure and it doesn't work. That's an example of the delta of decay. It's the distance between your last verified reality and the one that you're operating on right now. In most organizations, nobody knows how wide that gap really is, and nobody even really knows to watch for it or to look for it. The organizations that have navigated comparable failures that have caught these kind of problems before it became a catastrophe didn't necessarily have better technology or better plans. They didn't have faster people or a bigger budget. But what they did have was a different type of relationship with their own data. At some point someone in those organizations asked the question is what this system is telling us actually true and accurate information? Not just does the system work. Systems can work perfectly even if they're operating on false information. But is what it's telling us actually accurate? That question is the differentiating variable between success and failure, not technology and not resources. Actual verification of information. That's resilience, the mechanism underneath both failures and why it shows up in organizations that look nothing like Knight and Capital or Target Canada. Now let's move into advantage. How can this be an advantage to you and your organization? This is where I'm going to give you one thing to take into this week. And it's not a framework or a model or a reading list or anything like that. This is just a simple question. And it's something that's applicable today and that neither of these organizations ever took the time to ask. So here's the one thing. Identify one critical operations process in your organization. Find one where the information moves from one system or one team or one stage to the next, and where a downstream decision gets made based on what that information is and where it arrives. Now ask one question. At each handoff point in that process, is there anyone that's verifying that the information that arrived is accurate before action is actually taken? Not does the system work. Both of these systems in Knight Capital and Target Canada had functioning systems that appeared to be working. Not did the process get followed, because both these organizations followed their processes. The question is, is the information that lands at each step verified, vetted, and confirmed as accurate before it becomes the basis for that action? That is the gap that neither of these organizations found in time. So pick one process, map the handoffs out, and ask where that verification lives or doesn't and figure out if it needs to be addressed. One more thing, and this matters practically. The verification mechanism that you find has to match the nature of where the risk lives. Let me say that again. The verification mechanism must match the nature of the location or where the risk lives. How do we find that out? Not through a tabletop exercise, a tabletop exercise would not have surfaced the night capital problem. The legacy code on that one server was invisible to any process level review. It only became visible when live transactions were actually executed. The verification mechanism that could have caught it was actually a controlled deployment test in a non production environment, a simulation that could have shown live marketing conditions before the market opened and tested it to see if it was working properly. In the example of Target Canada's data corruption, the data corruption would not have been visible in a process review either. It was actually in the foundation of the data itself. The verification mechanism that could have caught this would have been a live parallel run, something that populated the ERP with real inventory data and testing replenishment to see if it was working and chucked against the actual physical stock before going live. Things like stress testing, what we call in emergency management red teaming, having parallel runs or live fire type exercises, these things aren't bureaucratic compliance activities that we have to jump through the hoops or must do to check the box. These are mechanisms that can help find the gaps in the space between your steps before your live system does or before a crisis shows that gap. The question is whether the practice is designed to surface what's actually there or to confirm that the documented processes look correct on paper. That's the true difference. In emergency management training, we use a concept called the red slice to explain why the verification matters and why the absence of it is so dangerous. Think of your organization's information picture as a pie. The green slice is what you know that you know. This is confirmed information. It's been verified, vetted, and it's actionable. You can build your decisions on the green slice. The yellow slice is what you know that you don't know. You've identified a gap or you have a question, there's something that's still unconfirmed, but you're aware of it. This yellow slice is manageable because you can see it. You can vet it and address it and resolve it before it causes damage downstream. Then there's the red slice. This is what you don't know that you don't know. This is the most dangerous ground in any organization. You're not looking for it because you don't know it exists. Decisions get made and actions get taken, systems are activated, all without any awareness that a critical piece of the picture may be missing entirely. In the case of Night Capital, their red slice was a single server that ran on a legacy code that nobody knew was still active. In the case of Target Canada, their red slice was an entire ERP foundation that was built on data that had never been confirmed against physical reality. So the goal, both in emergency management and in your organization, is to keep that red slice as small as possible at all time. Stress testing, parallel runs, and red teamings are the mechanisms that pull things out of the red and into the yellow. They identify these things that you don't know that you don't know. For you can see them now, you can address them, and you can take action to resolve them before they go live or before the crisis hits. That's information hygiene and information management in operational terms. And that's the discipline that neither of these organizations had in place when it mattered the most. So what happens if you do find a gap? If you find a gap and there's a real chance that you will, here's how to think about what to do with it. First thing to do is recognize the gap. Name that information and operational reality that have separated. Don't minimize it and don't escalate it before you understand it, but recognize it for what it is. Then identify which system is the source. Is this a pathway failure? The information exists, but it can't reach the people who need it. Or is it a foundation failure where the data itself is wrong? If you've recognized it and you've identified the problem, now stabilize it. Before you make an operational decision downstream, establish what's actually true and verify your actions before you take action, not after. Then evaluate what happened. Was the corrective action that you took or were about to take based on the real failure or a misdiagnosis of it? Night Capital's team failed here. Make sure that you don't. So if this episode surfaced a question that you don't have a quick answer to, about your data accuracy or about whether your dashboards are reflecting operational realities, or about whether a corrective action that your team runs regularly is actually solving the right problem, that's the right discomfort. It means you're looking at the right gap. The signal integrity assessment is a structured place to take that question. It's not an audit, it's a diagnostic, and it's built around whether your critical information pathways are actually functioning that the way that your organization assumes that they are. You can find the details at Southwindplanning dot com. That's advantageous one question, one conversation before the end of this week. Now let's bring this episode home. We had two organizations, two industries, both of them followed their processes, but both failed. Not at the decision points, but at the spaces in between those points. Where information was moved from one part of the system to the next, and where nobody verified that what arrived was accurate before it was acted on. The corrective actions made it worse because they were built on a picture that hadn't been confirmed, and the differentiating variable isn't better technology or faster response. It's verification of the information at the handoff before the information becomes basis for your actions. And the only way to find that gap is before your live system does, and to do that, you need to find a design, a practice that's built to surface what's actually there. In the next episode, we'll talk about another critical failure or critical incident that occurred that we can hopefully avoid in the future by taking the proper action. The blue to gray deep dive is paired to this episode and it goes deeper on both of these cases. The emergency management doctrine, how it's connected, the information hygiene framework and how that's connected, and what the red slice looks like when you map it against your own organization. If you're already a subscriber to the Blue to Gray newsletter, then it's already in your mailbox. If you're not a subscriber, you can subscribe for free at blue to gray.beehive.com. What the hell happened is produced by Southwind Planning Solutions LLC. If this episode was useful, the Blue to Gray newsletter goes deeper into this and many other topics each week. You can subscribe for free on Beehive. The link is in the show notes. If you're sitting on a subject that you've not verified anymore that you'd like to do. Or maybe there are other challenges where reality is interrupted in the planet. The starting point for that in the show notes. You can find us on Apple Podcasts, Spotify, or anywhere else to listen to the podcast. You can also find me on LinkedIn and on our website at www.selfwindplanning.com. I'm Michael Cracker. Thanks for listening, and I'll see you next time.