What the Hell Happened (and what can I do about it?)
Something happened this week that changed the conditions your organization is operating in. Most leaders will find out too late.
What the Hell Happened? is a weekly 30-minute briefing for executives in healthcare, higher education, manufacturing, logistics, transportation, and construction. Every episode takes the week's most consequential events — policy shifts, system failures, supply chain disruptions, regulatory changes — and works through what they mean for the people making decisions at the top.
Three segments. Every episode.
Readiness — what happened and why it matters. Resilience — what it reveals about the assumptions your organization is running on. Advantage — one specific action before Monday.
Hosted by Mike McCracken, founder of Southwind Planning Solutions, with decades at the intersection of emergency management and private sector operations.
Nothing about the Homeland Security Information Network ever looked broken. Alerts kept going out. Information kept flowing. That's what made it dangerous.
In this episode, Mike breaks down the HSIN breach — discovered in early July after weeks inside the platform federal, state, local, and private-sector partners use to share threat intelligence — and the failure underneath it that has nothing to do with cybersecurity expertise. Every agency pulling information off that platform had a source with a track record. What none of them had was a running answer to a second, separate question: is this specific information still accurate, right now — because the first answer had been "yes" for so long that nobody thought to keep asking the second one.
That distinction — trusting a source versus verifying today's claim from it — is the same discipline intelligence analysts are trained to hold onto, and it's the same discipline that failed inside SolarWinds in 2020, when a digitally signed, routinely trusted software update carried a backdoor into roughly 18,000 organizations, including several federal agencies, for months before anyone noticed.
This episode walks through:
What actually happened inside HSIN, and why the platform staying operational the entire time is the real warning sign
The difference between trusting a source and verifying what it's telling you right now — and why almost every organization quietly collapses the two into one
How the SolarWinds compromise shows the same pattern at a different scale
Why vetting a vendor once at the start of a relationship isn't the same as ever checking again
Four low-cost, low-effort ways to keep confirmed information from quietly turning into stale assumption — without adding headcount or new software: expiration dates on operational data, trigger-based audit gates tied to change management, tolerance-bandwidth alerts, and the ten-minute pre-mortem
A three-step, no-new-software exercise for finding the sources your organization has stopped questioning, before something forces the question for you
Mentioned in this episode: Business Lifelines™, Information Hygiene, the Signal Integrity Assessment
www.southwindplanning.com
www.bluetogray.beehiiv.com
SPEAKER_00
You probably have at least one system in your business right now that you trust completely and never think to double check. There's probably a vendor portal that tells you that everything's on schedule, or a regulatory alert feed that tells you that nothing's wrong. It could be an industry status page that you glance at once in a while and you never think about again because it's always been right before. Every organization has these trusted sources. Every trusted source eventually changes. The dangerous part here isn't that they change, it's actually that the organizations rarely notice when these changes happen. Sometimes it's a policy shift. Sometimes it's a system that fails. Or it could be that a market moves. Often leaders find out too late. The signal was there, but nobody turned it into something they could act on before the moment passed. What the hell happened is a weekly briefing for leaders who would rather be ready. I'm Michael Crack. I've spent over 30 years working in this situation. Planning gets interrupted by reality. This program looks at not just what the hell happened, but also what can be done about it. Right now, national security agencies are trying to untangle a massive, weeks-long breach of their primary intelligence platform. This is a system that never once threw an error code while hackers actually lived inside of it. This is a masterclass in how trust fails, and if you think that your vendor portals and regulatory feeds are immune from this type of attack, stick around. Because what broke at the federal level is likely sitting inside of your business today. Hold on to that thought for a second. Back in late May, or maybe it was early June, the timeline's still a little bit fuzzy. The exact same kind of trust failed at the federal level. Somebody got inside of the Homeland Security Information Network, the platform that federal, state, local, and even private sector partners use to share threat intelligence, and they also use it to coordinate security. Somebody got inside this system and they stayed there for weeks. Nobody noticed until the first days of July, and this was about the time when the system was being leaned on heavily to help coordinate security for the World Cup games that were currently underway across the country. DHS has confirmed this breach, and the hackers got into the HSIN servers and also a connected SharePoint system. This hasn't been attributed to any specific country or group yet. The Department of Homeland Security says that they've isolated the affected systems and that classified networks weren't touched. That the platform stayed operational the entire time. That last part is the key. The system never actually had to go down for something to be terribly wrong. During a hearing of the Senate Intelligence Committee, one senator simply put it like this The information on the Homeland Security Information Network isn't classified, but it is highly sensitive, and losing control of it puts national security at risk. Now normally this is where I'd spend a few minutes translating a government story into something that you would recognize maybe in your own building, but not this time, because you already recognize it. If your organization has ever leaned on a regulatory feed or a vendor status page or an industry alert system to tell you whether something's safe, then you are already running some smaller version of exactly this type of system that just failed at the Department of Homeland Security. So the question that we're going to carry into the next part of this podcast is not could this happen to me? Because it already has quietly, on a smaller scale and you probably didn't even notice. What's the one external source that your business treats as automatically true? And when is the last time that you checked on that source to see whether the source itself was still trustworthy, not just the information that it was providing to you? Nothing about this platform ever really looked broken, and that's the strange part. There were alerts that kept coming out, and information kept flowing, even with the alerts. What broke was something that nobody was watching for. It was the assumption that the platform itself was still trustworthy in the first place. Every agency that was pulling an alert off of the Homeland Security Information Network during these weeks wasn't looking at bad information because someone had lied to them. They were actually looking at a source that they had every reason historically to trust completely, but the source just completely and quietly stopped earning that trust, and nobody had rechecked that because nobody thought they needed to. This is not a data problem. This is more of a trust maintenance problem. And trust, once you've got it established, almost never gets re-examined on its own, and that's the problem. It just carries forward week after week until something forces the question or forces it to be challenged. Here's something that's interesting. People who spent decades studying intelligence have learned something that most organizations don't often teach to their employees or their trainees. They discovered that every piece of information requires two completely different judgments. The first one asks, can I trust this source? And the second one asks, is this particular information actually true? These may sound like the same question, but they're really not. Most organizations can answer the first question and they only have to answer it once, and then they spend years assuming that they've answered the second one as well. Here's the part that gets hard to separate and often never gets taught or trained on. When you decide to trust something, you're actually answering two different questions that I just mentioned before, and almost everybody collapses them into one single question. The first one asks about the source, about whether this information or this thing has been reliable before and looks at its track record to see if it's earned a benefit of the doubt. But it's that second question about what's in front of you right now or today, and whether or not that that specific piece of information is true regardless of who's handing it to you, where it's coming from. There are intelligence analysts that are trained to keep these two questions separate on purpose. Their training exists precisely because people have got them tangled without having this type of training. A source with a spotless history can still hand you something false, and a source that you've never even heard of, one that you'd have every reason to doubt, can occasionally hand you something that turns out to be exactly right or exactly accurate. The moment that you let a source's reputation answer the second question for you, you've stopped evaluating anything, and now you're just deferring information. That's the failure that's sitting beneath this problem with the Homeland Security Network. Every agency that was pulling information from that platform had a source with a track record. What they didn't have though was a running answer to the second question. Is this information or this specific thing right now still accurate? Because the first answer had been yes for so long that nobody thought to keep asking the second question. This isn't the first time that a trusted channel like this has carried something that it shouldn't have. Back in December of 2020 there was a company called Solar Wind, and they discovered that a piece of network monitoring software that was used by roughly 18,000 organizations that even included several federal agencies had been shipping with a backdoor that was built into it for months. The update was digitally signed and it came through the normal channels the way it always had. The companies that were running it weren't careless, they did exactly what all of us do with a vendor that we've used for years. They let the relationship stand in for the check. That's the pattern from both of the stories and the one that they share. This isn't one bad decision, it's actually a good decision that was made once, but then it quietly stopped getting remade. Here's some other examples. Think about your CPA or your outside attorney. Maybe it's your payroll provider or your HR platform. Could even be your bank or something like that, but nobody re-verifies these sources of information. Not because they're bad, but it's because information like this and these sources has been something that's earned your trust over time. And that's universal. There's an organizational reason that this keeps happening, and it's not that anyone's being careless. It's that most vendors and sources have vetting that happens once within an organization at the beginning of that relationship. This is when everyone is paying really close attention because the relationship is new and unproven. There are procurement that runs checks on the purchasing, there are legal reviews to the contract, and there are people who confirm the credentials. And then because that first pass was thorough, this source now gets treated as permanent and accepted. Nobody ever schedules a second look because scheduling a second look on something that's never failed feels like it's solving a problem that you don't really have. Feels like a waste of time. Sometimes I see a version of the same mistake and the responses that I have made on emergency management situations, and it's one of the first things that gets corrected early on in our area. Things like mutual aid agreements, resource list, maybe a contact roster for who's available during a disaster, none of that gets vetted only once and then filed away in a drawer somewhere. That information has to be revalidated on some sort of a cycle, either every planning period or every couple of months when the plan's reviewed, whether or not anything's gone wrong, because the alternative is to discover that there has been an error during an actual emergency and that the resource that we were counting on quietly stopped being available several months before possibly, and that nobody updated the list or checked to verify that it was still valid. The discipline part of this all isn't really distrust. It's just refusing to let we check this already be the answer for we checked this recently. That's the habit that the Homeland Security Network was missing, and that was at a scale that most of us will never operate at. It's also very likely the same exact habit that could be missing in somewhere in your organization on a much smaller scale right now. There's a name for this in the work that I do with private industry companies. I call it information hygiene, and the short version of this is that the information exists in different states. There's information that sometimes has been confirmed by a source that you trust, and some of it that you simply trust and assume is still true because it always has been. But most of the failures that I see or that I look at don't come from information that's gone missing. It comes from information or something that's been confirmed quietly and then slides into something that's merely assumed, with nobody noticing the slide. Think about what this might look like. I mentioned earlier that it could be a regulatory feed or a supplier portal, or maybe it's an industry alert system. You didn't do anything wrong in getting it set up, and that's what makes this so hard to catch. There's really no bad decision to point at. The failure here is in the ordinary human tendency to stop questioning something the longer that it's been a reliable source. The longer that a source has been right, the less likely that anyone will be to ask whether it's still the same source as it was when you first decided to trust it. So what does this all mean? It means that organizations that catch this before it costs them anything aren't running better technology. They've built an actual habit, they've put it on a calendar with a name attached to it, of going back and reverifying the sources that they've already decided to trust. Not the information coming out of the source, but the source itself. That's a different discipline than most risk programs are built for, and it's the one that this episode is really about. One thing worth saying directly because some of you are probably thinking it. I've spent thirty years in emergency management. That's emergency management, not cybersecurity. And I just walked you through a federal network breach, a software supply chain compromise, at least not in the way that people inside a security operations center or the IT center of these companies would know it. But here's what I do know, and this is the only thing that this show has ever actually been trying to sell. It's the pattern that exists beneath these type of systems. It's a trusted source or a track record that's standing in for a current check when nobody is there to re-verify because nobody thought they needed to. So it's not really a cybersecurity pattern. I've watched that exact shape show up in evacuation plans, mutual aid agreements, resource lists that look fine on paper, and then they fell apart the moment that somebody actually needed to use them. I spent years watching that type of failure repeat itself across situations with nothing else in common, and that's what's let me to walk into a story about a federal network that I've never heard of a few weeks ago and then tell you with real confidence what went wrong with it. This is also probably why you would consider bringing someone in from the outside rather than someone from the inside who knows your specific system. You already have the people who understand your vendor feeds or your supplier portals or your regulatory dashboards better than anyone outside ever will. But what most organizations don't have is someone that's trained specifically to see the shape of a trust failure before it becomes a headline. That's not really industry knowledge, it's more of a pattern recognition knowledge and understanding. And it travels into many rooms or operations center that I've never worked in before because what actually breaks is never really the technology. It's always the same type of habits that are missing in the same type of places. Most leaders who hear this story will walk away quietly and add another layer of monitoring to their systems. They'll add another dashboard or another alert that nobody has time to look at or read. None of that would have caught what happened at the Homeland Security Network. Preventing this assumption drift without burning a budget or creating a dedicated data watcher role boils down to automating the pulse checks and building a validation trigger into your existing workflows. If a piece of information rarely changes, you don't need to put a human being there to sit and stare at it. You just need to build a structural system that forces the information to verify itself when it actually matters. Let me give you four examples. These are low cost, low effort strategies that can help keep your confirmed data from turning into stale assumptions. First, set expiration dates on operational data. In IT and computer networking, data has a time to live, an automation clock that tags data with a hard expiration date. When the time to live period expires, the system doesn't necessarily delete the data but it simply flags it as now unverified, and it needs to be reconfirmed by the user before it can be used. Here's how to apply this without using dedicated staff. Instead of monitoring the data continuously, set an automated calendar or an alert that's tied to the data's risk level. For example, if it's a critical safety spec, it could be every six months. But if it's something more routine like policy documents, maybe it's not that regular, maybe it's yearly, or even every two years. The point being that it needs to be checked and set it to an automated calendar so it doesn't get forgotten. If there's a team member that relies on a certain piece of data that pops up then as marked unverified, they would be required to perform some kind of a check, a 30-second sanity check, or just a verification before they can act on it. So the data isn't monitored daily, but it is revalidated at the point of use. Another idea is to set trigger-based audit gates. Instead of checking data on a set timetable, you could tie the data re-verification to changes in the surrounding environment. The data you would usually degrade because something else shifted around it, either a system update or a vendor change, or maybe a tweak to one of your processes. So again, to apply this without dedicated staff, simply add a standard line item to an existing change management checklist. In other words, if we change step A, what legacy assumptions in step B or step C are we relying on that may no longer be true? Another idea would be tolerance bandwidths. If the data stays within normal operational parameters 99% of the time, then monitoring it manually is a waste of money. Instead, one option would be to set up automated thresholds or exception reporting that would only alert when a variable steps outside of historical baselines. To implement this suggestion without dedicated staff, simply use basic software automation tools like Excel scripts or system alerts that remain silent until a boundary is crossed. Then you ignore the data when it's behaving normally, but the moment that it drifts into an edge case or something on the fringe of what's acceptable, an automatic flag forces a pause. There are also things like kill switches or pre-mortems that you can use or that you can do to help minimize the risk. Changing the culture in your environment or your organization costs nothing to make adjustments to. One of the primary reasons that stale assumptions survive is because of confirmation bias. This is where people assume that just because something worked yesterday, it must still work today. So the next time before you launch a major project or execute a critical process, think of this. Run a 10-minute pre-mortem. Ask your team to assume that the project if the project fails catastrophically six months from now, what core assumptions did they make or did they take for granted that turned out to be wrong? This will force the team members to pull out old confirmed facts and ask, is this actually still true or are we just relying on memory? What actually catches this is smaller than what you would think, and it takes less time than the meeting that you probably sat through this morning. It's three simple steps. First, list every external source that your business treats as automatically true without a verification habit that's attached to it. Things like regulatory feeds or vendus vendor status pages or industry alert systems. Anything where you receive information and act on it because the source has always been good and accurate, not just because you checked it recently. Next, for each one of these, attach a name to it. Identify one person who's specifically responsible for that particular source, not just the information it delivers, but responsible to make sure that the source itself is verified and accurate. And finally, put a date on your calendar. Not a date that says we'll check when something feels off. An actual recurring date. Quarterly is probably going to be plenty, where a person re-verifies that the channel they're using is still what it was the day that they started using it. That's the entire exercise. There's no new software, no new headcount, no new vendor relationship. It's just a little bit of time to make a list, attach a name to it, and then pick a date and how often that that needs to be verified. Here's what it looks like depending on where you sit. In the healthcare vertical, it could be a regulatory or a public health alert feed that your team leans on the hardest. Find out who's responsible for confirming that the feed hasn't been altered upstream somewhere. And if the answer is nobody, then that's the first name to add. Maybe you work in manufacturing and it's a supplier status or an industry disruption alert. Someone needs to own reverifying that channel on a schedule, not just consuming whatever it says. Maybe you work in construction or infrastructure and it's a safety regulation or a notification system driving your operational decisions. Sometimes these could even be life safety decisions that nobody's ever asked is still the same channel as it was the day that you signed up for it. These are exactly the kind of blind spots that things like the Signal integrity assessment is built to surface. This is not to identify or say that your systems are broken because they're probably fine. But it does help to identify where nobody has specifically been looking at the difference between the information you've verified and the information that you've just got used to trusting. If this episode has made you a little bit uncomfortable about a source that you've never rechecked, then that discomfort is the whole point. Hopefully, this doesn't actually feel like a cyber episode because it wasn't intended to be that. Hopefully, it feels like an episode about maintenance. What I mean here is not equipment maintenance, but decision maintenance and trust maintenance and information maintenance. That's going to fit almost perfectly with how we can work on our resilience. So here's a quick recap of what we talked about in this episode. What the hell happened? A trusted government platform got quietly compromised for weeks and nobody noticed because the platform never stopped looking normal. What did this tell us? The most dangerous information failures aren't the ones where the data goes missing. They're the ones where a source that you've always trusted stops getting checked, and then nobody notices it the moment that it happens. What the hell can I do about it? You can list the sources that you treat as automatically true and name an owner for each one of these sources, and put a real date on when that source should be re-verified, not just the data that it's telling you. Remember that most failures that I study don't come from information that goes missing. The failures come from something that's confirmed and then quietly slides into something that becomes an assumption with nobody noticing that slide. Next week we'll try to stay close to the same territory because it's important. So if this one got under your skin a little, stick around and see how that's about. If you want the deeper version of this argument or discussion, check out the blue-to-gray newsletter that pairs with every episode. And this week's issue will go further on the information hygiene framework that we've talked about. And it'll also point to where we're headed next. I'm Mike McCracken. This is What the Hell Happened, and I'll talk to you next time. If you enjoyed this episode, drop me a comment or an email. It's in the show notes. I'd appreciate hearing from you. I'd also appreciate hearing any suggestions that you have for new ideas or new content. I'm always looking for things that will be exciting and interesting to listen to. What the Hell Happened is produced by Southwind Planning Solutions LLC. If this episode was useful, the Blue to Gray newsletter goes deeper into this and many other topics each week. You can subscribe for free on Beehive. The link is in the show notes. If you are sitting on an assumption that you have not verified in longer than you'd like to admit, or maybe there are other challenges where reality has interrupted your planning. There's a starting point for that in the show notes as well. You can find us on Apple Podcasts, Spotify, or anywhere else at ToListen Podcasts. You can also find me on LinkedIn and on our website at www.southwindplanning.com. I'm Mike McCracken. Thanks for listening, and I'll see you next time.