All Foundry Insights

Foundry Insight

What a Broken Cooler Taught Me About Business Continuity

Organizations are designed around the assumption that work will proceed as planned. It rarely does. When a routine equipment failure disrupted a retail operation, the people doing the work adapted—and revealed something important about how resilient systems actually function.

It wasn’t a cyberattack. It wasn’t a hurricane. It wasn’t a ransomware incident that brought an organization to its knees. No executive emergency meeting was called. Nobody declared a disaster. There were no flashing dashboards, no crisis communications team, no frantic conference calls with vendors trying to restore critical infrastructure before the morning news cycle began. In fact, if you had walked into the building that morning, you might not have noticed anything unusual at all.

A piece of equipment had failed. That was it. Fifteen years ago, that’s probably all I would have seen. A broken refrigerator. An inconvenience. A maintenance ticket. Something for Facilities to deal with.

Over the years, however, I’ve come to realize that organizations are constantly telling us who they really are. Most of the time, we’re simply too busy looking at the symptom to notice the system behind it.

Text Box: OBSERVATIONDocumentation describes an organization.Relationships operate it.Experience has changed the questions I ask. I no longer find myself asking, “What broke?” Instead, I ask, “What just became visible?” Because every disruption, no matter how ordinary, reveals something about the organization experiencing it. It reveals where communication really happens. It reveals whether authority is centralized or distributed. It reveals how people respond when the process no longer matches reality. It reveals whether teams protect their own priorities or instinctively collaborate toward a common mission. Most importantly, it reveals whether an organization has built the capability to adapt, or whether it has simply built the illusion of control.

For decades, organizations have invested enormous effort into becoming predictable. We write policies, document procedures, define workflows, implement governance frameworks, and purchase increasingly sophisticated technology to reduce uncertainty. Those investments matter. Good governance, sound engineering, and disciplined operations are not bureaucratic exercises; they are the foundations upon which resilient organizations are built.

But there is an uncomfortable truth that every experienced operator eventually learns. Reality does not read our documentation. Reality has no interest in organizational charts. Reality does not care how elegantly we’ve diagrammed a workflow or how thoroughly we’ve cataloged our risks.

Reality simply asks one question: Now what?

And it asks that question over and over again.

A supplier misses a shipment. A key engineer resigns. A network link fails. A software deployment introduces an unexpected bug. A storm knocks out power. A customer changes their demand overnight. Or sometimes…a refrigerator stops working.

When those moments arrive, something remarkable happens. The organization’s formal architecture, the one captured in policies, process maps, and PowerPoint presentations, quietly steps into the background. Another architecture emerges. An invisible one.

Text Box: Figure 1 - Hidden Operational SystemIt exists in relationships rather than reporting structures. In trust rather than authority. In judgment rather than procedure. In the countless unwritten decisions that ordinary people make when no document can tell them exactly what to do next. I’ve become convinced that this invisible architecture is the real operating system of every organization. Most days, we barely notice it.

Then something breaks. And suddenly, everything depends on it.

Recently, I had the opportunity to watch that happen in one of the most unlikely places imaginable. Not inside a Security Operations Center. Not during a disaster recovery exercise. Not at the headquarters of a Fortune 500 company.

I watched it unfold in the Digital grocery operation of a Walmart store after the department’s walk-in cooler failed. At first glance, it looked like a straightforward facilities problem. Within a few hours, I realized I wasn’t watching a broken cooler. I was watching an organization reveal itself.

Hidden Architecture

By the time the cooler had been emptied and refrigerated inventory relocated to the dairy department, the immediate crisis had passed. The products were safe. Orders could still be fulfilled. Customers could still receive groceries. If you had been evaluating the situation from thirty thousand feet, you might have concluded that the business continuity plan—formal or informal—had worked exactly as intended. A critical capability had been restored. Mission accomplished. Except that’s not what I was watching. What fascinated me wasn’t the workaround itself. It was everything the workaround revealed.

One of the most persistent misconceptions in organizational design is the belief that organizations operate according to the diagrams we create. We draw organizational charts showing who reports to whom. We publish process maps illustrating how work is supposed to flow. We document policies, procedures, escalation paths, approval chains, and governance models until every responsibility appears neatly assigned. Those artifacts are valuable. They create consistency. They establish accountability. They communicate intent. But they are not the organization. They are a description of the organization we hope exists. The real organization lives somewhere else. It lives in the conversations that happen between departments before anyone opens a ticket. It lives in the trust that develops between people who have solved problems together for years. It lives in the experienced associate who quietly notices a problem before anyone else does and starts fixing it without waiting for permission. It lives in the supervisor who understands that protecting the mission sometimes requires bending yesterday’s workflow to preserve today’s outcome. It lives in the countless small judgments that no policy manual could ever anticipate.

I have come to think of these as two entirely different architectures:

  • The first is the formal architecture.
  • The second is the operational architecture.

Text Box: Figure 2 - Formal Architecture vs. Operational ArchitectureThe formal architecture is easy to document. It consists of reporting structures, documented procedures, governance frameworks, software systems, policies, and official responsibilities. Auditors can review it. Consultants can diagram it. Executives can approve it. It’s important. But it’s only half the story.

The operational architecture is considerably harder to see. It’s built from relationships rather than reporting structures. From judgment rather than procedure. From credibility rather than authority. From communication patterns that have evolved over months or years of working together. It is the invisible network through which work gets accomplished. Most of the time, these two architectures overlap reasonably well.

Processes function. Departments cooperate. Information flows where it’s supposed to go. The invisible architecture remains largely invisible. Then something breaks. That’s when the distinction becomes impossible to ignore.

The cooler failure didn’t simply create a logistics problem. It temporarily erased a portion of the store’s formal operating model. The documented workflow assumed refrigerated inventory would be staged in one location. Reality had other ideas.

Almost instantly, the organization began constructing an entirely new operational model. Nobody gathered in a conference room to redesign the process. Nobody updated a Standard Operating Procedure. Nobody waited for next week’s continuous improvement meeting. Instead, the organization began doing something that healthy systems do remarkably well.

It started experimenting. One rack moved here. Another associate covered that responsibility. Someone discovered that calling ahead saved a trip. Another person realized certain delivery windows created unnecessary congestion. A dispenser learned to anticipate what dairy would need before anyone asked. None of these adjustments were individually significant. Collectively, they were the new process being born.

That’s one of the remarkable things about resilient organizations. They rarely redesign themselves through a single brilliant decision. They evolve through hundreds of tiny corrections. Watching this unfold reminded me of something software engineers have understood for decades. When a distributed system experiences a failure, the interesting question isn’t whether one server went offline. The interesting question is how the rest of the system responds.

Does traffic reroute? Do redundant services take over? Does the platform gracefully degrade? Or does the entire application collapse because one component disappeared?

Organizations behave the same way. The cooler wasn’t the system. The cooler was one component within a much larger operational network. The real system consisted of people. Communication. Trust. Judgment. Timing. Coordination. Shared purpose. The equipment failure merely illuminated those connections.

In retrospect, the most important thing that failed that week wasn’t refrigeration. It was an assumption. The assumption that the documented process was the process. It wasn’t. It never is.

Every organization has an invisible operating system running beneath its policies and procedures. Most leaders never see it because, under normal circumstances, they don’t need to. Work gets done. Customers are served. Objectives are met. The invisible architecture quietly does its job. Until reality changes.

That’s when organizations discover whether they’ve built a culture capable of adaptation—or merely a process capable of repetition. And that, I think, is one of the most important distinctions in organizational maturity.

Every Solution Creates New Problems

One of the most counterintuitive lessons in systems thinking is that solving a problem and moving a problem are rarely the same thing. The distinction matters because organizations are exceptionally good at celebrating the first while overlooking the second. The cooler failure was resolved quickly enough. Refrigerated inventory was preserved, customer orders continued to flow, and the immediate threat to the business was contained. From a traditional incident management perspective, the response looked successful. Yet every decision that stabilized one part of the operation quietly introduced new demands somewhere else.

Moving refrigerated inventory into the dairy cooler solved the storage problem, but it lengthened every picker’s route through the store. Longer routes meant more time between selections. More time meant fewer completed orders per hour. Fewer completed orders increased pressure on staging. Staging delays affected dispensers waiting outside with customers. Every additional minute spent walking across the building became a minute unavailable for serving someone else. None of these changes were catastrophic in isolation. Together, they formed an entirely new operating environment.

That is how complexity usually behaves. It rarely disappears. It changes shape. Organizations often describe this as a tradeoff, but I have come to think of it differently. Complexity is remarkably faithful. If you force it out of one corner of the system, it simply finds another place to live. We don’t eliminate complexity nearly as often as we relocate it.

Text Box: Figure 3 - Complexity RelocatesThe relocation of inventory also changed communication patterns. What had once been a short conversation between teammates working in the same physical area became coordination between departments separated by the length of a supercenter. At first there were no radios available to bridge that distance. Instead, associates became the communication channel. Someone would walk to dairy to ask for an item. Another person would return with an answer. A third might discover that conditions had already changed. Every trip consumed time, but more importantly, every trip represented information that the original workflow had assumed would travel almost effortlessly.

Watching those extra steps accumulate was strangely fascinating. It reminded me that organizations spend enormous effort measuring visible work while invisible work often goes uncounted. We measure the order picked, the delivery completed, the ticket closed, or the feature deployed. We rarely measure the conversations required to make those things possible, the context switching, the interruptions, the walking, the waiting, or the countless micro-decisions that fill the spaces between documented tasks. Yet those invisible activities frequently determine whether a system feels effortless or exhausting.

As the days passed, something encouraging happened. Associates began anticipating the friction before it occurred. Pickers grouped requests together. Dispensers learned when it was worth making a single trip instead of several. Dairy associates anticipated what Digital would need during peak periods. No one issued a new procedure. No committee approved a redesigned workflow. The organization simply learned. The relocated complexity had not vanished, but people were steadily reducing its cost through experience.

This is why I have become skeptical whenever someone claims that a process has been ‘optimized.’ Optimized for what? Optimized under which assumptions? Every process is optimized for the environment in which it was created. Change the environment and yesterday’s optimization often becomes today’s constraint. Mature organizations understand this instinctively. They treat processes as living hypotheses rather than permanent truths, refining them whenever reality exposes a better path.

Looking back, the most valuable observation from that week was not that people worked harder, although they certainly did. It was that they continually searched for ways to make the work easier for one another. Every improvement, however small, reduced friction somewhere in the system. That is what operational maturity looks like in practice. It is not the elimination of problems. It is the disciplined reduction of unnecessary friction while the mission continues uninterrupted.

Adaptation Isn’t Elegant

By the second day, the novelty had worn off. Whatever optimism accompanied the initial response had given way to the reality of sustaining it. Temporary work has a way of revealing itself over time. What seems entirely manageable for an hour begins to feel cumbersome after an entire shift and exhausting after several days. The organization had solved the immediate problem. Now it had to live with the consequences of its own solution.

Text Box: PATTERNComplexity is rarely eliminated.It is relocated.From a distance, adaptation often appears clean and deliberate. We speak about organizations “pivoting” or “adjusting course” as though change happens in a single, coordinated motion. The reality is considerably messier. Adaptation is usually improvised. It is built from small observations, imperfect experiments, and countless conversations between people trying to make the next hour work just a little better than the last one.

One of the clearest examples was communication. Under normal conditions, the Digital team and the dairy department rarely needed to coordinate at this level of intensity. Once refrigerated inventory was relocated, however, every order depended on information moving between two groups separated by the length of the building. There were no purpose-built communication channels waiting to be activated. Associates became the network. They walked requests from one department to another, confirmed inventory, answered questions, and carried updates back across the store. It was inefficient by almost any engineering standard, but it was sufficient to keep the operation moving.

What struck me was not that the process was inefficient. It was that everyone recognized it was inefficient. Nobody mistook the workaround for the destination. Each trip across the store highlighted another opportunity to reduce unnecessary movement, eliminate an extra question, or anticipate the next request before it had to be asked. The workaround was never accepted as “good enough.” It became a living prototype that people refined every time they used it.

That distinction matters because mature organizations rarely confuse adaptation with optimization. Adapting means accepting that today’s solution is temporary and remaining willing to improve it tomorrow. Optimization, by contrast, is the result of many successful adaptations accumulated over time. One is a behavior. The other is an outcome. Organizations that skip directly to optimization often stop learning just when learning matters most.

As the week progressed, patterns began to emerge. Associates naturally synchronized their work. Requests were bundled together. Peak periods became more predictable. People anticipated one another’s needs without being asked. None of these improvements appeared in a revised process document, yet every one of them reduced friction. The organization was slowly constructing a better operating model from experience rather than instruction.

I often think governance is misunderstood in moments like these. Good governance is not the rigid enforcement of yesterday’s process despite today’s reality. Nor is it permission for everyone to invent their own way of working. Good governance provides a stable purpose, clear boundaries, and shared accountability while leaving room for informed judgment inside those boundaries. In other words, governance should create the conditions in which intelligent adaptation can occur safely.

Looking back, I no longer remember every individual adjustment that was made during those days. I remember something more important. I remember watching ordinary people solve ordinary problems with remarkable consistency—not because they possessed the perfect process, but because they shared a common objective and trusted one another enough to keep improving the work together. That is adaptation in its truest form. It is rarely elegant, seldom visible, and almost never celebrated. Yet it is the quiet discipline upon which resilient organizations are built.

The Customer Only Sees the Surface

Several days into the disruption, a customer asked a simple question: “Are you out of milk?” It was an entirely reasonable question. From their perspective, the online ordering system no longer offered the selection they expected. The most obvious explanation was that the store had simply run out of product. The dairy cooler was full. Milk was available. What had changed was not inventory, but operational capacity.

Text Box: QUESTIONWhat invisible work keeps this process alive?Early in the incident, store leadership made a deliberate decision to limit the quantity and variety of refrigerated items customers could order online. It was a pragmatic choice. Every refrigerated order required additional travel, additional coordination, and additional labor. Left unmanaged, demand could have quickly outpaced the team’s ability to fulfill orders accurately and on time. Rather than allow service quality to deteriorate across the board, the organization intentionally constrained one part of the customer experience to preserve the whole.

That decision illustrates one of the quiet realities of operational leadership: customers rarely experience the problem itself. They experience the decisions organizations make in response to the problem. A failed piece of equipment is invisible. A delayed shipment is invisible. An exhausted associate walking an extra mile through a store is invisible. What customers see are empty menu options, longer wait times, unexpected substitutions, or a product that is suddenly unavailable. They encounter symptoms while the organization wrestles with systems.

This disconnect exists in nearly every industry. A software customer sees a feature that loads slowly without realizing an engineering team is balancing traffic after a failed server. A patient experiences a delayed appointment without seeing the staffing shortage unfolding behind the scenes. A manufacturer misses a delivery date while dozens of people work tirelessly to recover from a supplier failure. In each case, the customer’s experience is real, but it represents only the visible edge of a much larger operational story.

It is tempting to think that the goal of resilience is to prevent customers from noticing that anything has gone wrong. Sometimes that is possible. More often, it is not. Mature organizations recognize that preserving trust is often more important than preserving the appearance of perfection. They make deliberate choices about where to absorb disruption, understanding that every system has finite capacity. The question is not whether compromise will occur, but where it can be done with the least harm to the mission and to the people the organization serves.

That week, the reduced selection of refrigerated items was not evidence of failure. It was evidence of prioritization. The store chose to narrow the promise it was making rather than overextend itself and fail to keep it. That distinction may seem subtle, but it reflects disciplined operational judgment. The easiest promise to make is an unlimited one. The harder—and often wiser—choice is to acknowledge reality and shape expectations around what the organization can reliably deliver.

The customer who asked about milk almost certainly left with the impression that the store was temporarily short on inventory. I don’t blame them. Without standing behind the scenes, there was no reason they should have reached any other conclusion. Yet that brief interaction stayed with me because it captured something much larger than a grocery order. It reminded me that perception and reality often diverge inside complex systems. The work of leadership is not merely to manage the system itself, but to understand how its invisible decisions become visible experiences.

The longer I reflected on that exchange, the more I realized it applies far beyond retail. Every organization leaves fingerprints on the experiences of its customers, employees, partners, and stakeholders. Those fingerprints are formed long before anyone answers a support call, delivers a product, or approves a document. They are formed in the countless operational decisions that shape how work is performed when conditions are less than ideal. By the time the customer notices the surface, the deeper story has already been written.

What Operational Maturity Actually Looks Like

By the time the cooler had become part of the background rather than the headline, I found myself thinking less about refrigeration and more about maturity. Not organizational maturity as it appears in assessment models or certification checklists, but maturity as it reveals itself under pressure. It is easy to mistake maturity for perfection because mature organizations often make difficult situations look routine. Yet what I witnessed that week was anything but perfect. There were extra steps, longer walks, improvised communication paths, and countless small inefficiencies. None of those contradict operational maturity. In many ways, they demonstrated it.

We often evaluate organizations by the artifacts they produce: policies, standards, process maps, dashboards, audit results, and performance metrics. Those things are important because they represent intentional design. They establish expectations, define responsibilities, and create a common language for the enterprise. But artifacts are static. Maturity is dynamic. It emerges in the moment when the documented process is no longer sufficient and people must rely on judgment, trust, and shared purpose to bridge the gap between what was planned and what reality demands.

This is why so many disciplines that appear unrelated gradually converge on similar ideas. Lean teaches us to remove unnecessary friction from the flow of work. High Reliability Organizations emphasize preoccupation with failure, deference to expertise, and continuous learning. DevOps encourages rapid feedback, collaboration across traditional boundaries, and relentless improvement. Mission Command asks leaders to communicate intent clearly enough that people can exercise initiative without waiting for permission. Although these approaches developed in different industries and for different purposes, they all recognize the same underlying truth: resilient performance depends less on rigid control than on enabling intelligent adaptation.

That realization changed the way I think about governance. Good governance should never become a substitute for thinking. Its purpose is to provide direction, establish boundaries, clarify accountability, and preserve organizational intent. Within those boundaries, however, people must still be free to observe, learn, and adapt. A governance model that cannot tolerate thoughtful adaptation is not creating resilience; it is merely postponing failure until circumstances exceed the assumptions upon which the process was designed.

As I reflected on the week, I realized that nobody had paused to debate whose department owned the problem. Facilities addressed the equipment. Dairy protected the product. Digital continued serving customers. Leadership balanced demand against available capacity. Every group contributed according to its strengths while remaining focused on a common outcome. That alignment was not accidental. It reflected a shared understanding that the mission was larger than any individual function. The cooler may have belonged to one department, but continuity belonged to everyone.

Perhaps that is the clearest indicator of operational maturity. It is not measured by how rarely problems occur, because every complex organization encounters unexpected disruption. It is measured by how quickly people move beyond the question of blame and toward the work of understanding, adapting, and improving. Mature organizations recognize that every disruption contains information. Every workaround teaches something about the system. Every recovery leaves behind an opportunity to become slightly more capable than before.

When viewed through that lens, the broken cooler was never simply an operational inconvenience. It became an unplanned exercise in organizational learning. The equipment eventually returned to service, but that was almost beside the point. The more enduring outcome was the knowledge the organization accumulated while the cooler was unavailable. Procedures may return to normal. People do not. They carry forward what they have learned, and if the organization is healthy, that learning becomes part of its capacity to meet whatever reality asks next.

Returning to Normal

Most stories about disruption end when the system is restored. The equipment is repaired, the lights come back on, the server is rebuilt, or, in this case, the cooler begins humming again. It is a satisfying ending because it suggests a return to equilibrium. The problem has been solved. Operations can resume. The organization can move on.

Reality, however, has a way of lingering.

When the cooler returned to service, refrigerated inventory moved back into its familiar location. The extra trips across the store disappeared. Communication between departments became simpler. The temporary workflow that had sustained the operation for days was no longer necessary. On paper, the Digital department had regained its full operational capacity overnight.

Text Box: Figure 4 - Recovery Has InertiaThe people had not.

What surprised me most in the days that followed was not that the team remembered how to perform the work. Of course they did. This was normal operation. Everyone understood the process. Everyone knew where products belonged and how orders were supposed to flow. Yet, despite the restoration of the physical system, the operational rhythm remained elusive. Four days later, the department was still struggling to achieve the same level of efficiency it had demonstrated before the failure.

At first, that seemed counterintuitive. If capacity had returned, why hadn’t performance?

The more I watched, the less surprising it became. Human systems adapt just as physical systems do, and adaptation leaves a residue. During the disruption, people had built new habits, new communication patterns, and new expectations. They had learned to compensate for constraints that no longer existed. Returning to yesterday’s workflow required another period of adjustment. The environment had changed again, and once again people were learning.

That observation has stayed with me because it challenges a quiet assumption that appears in many organizations: we tend to think of recovery as a switch. We imagine that once the technical problem is resolved, performance should immediately return to its previous state. But organizations are not machines. They are communities of people. Restoring equipment is often far easier than restoring rhythm.

Text Box: REFLECTIONRecovery has inertia.There was another asymmetry that fascinated me. Customers returned to normal almost immediately. As soon as refrigerated items became available again, order volumes recovered to their previous levels. Demand did not hesitate. The organization, however, needed time to absorb that demand with the same confidence and efficiency it had displayed before the disruption. Providers recovered more slowly than the people they served.

That distinction feels important because it reminds us that operational capacity and human capacity are related, but they are not identical. An organization can restore its infrastructure in a single afternoon while its people continue recalibrating for days. Leaders who fail to recognize that gap may conclude that the work is finished when one of the most important phases has only just begun.

In hindsight, the cooler failure did not teach me that resilient organizations always recover quickly. It taught me something both simpler and more profound. Mature organizations understand that recovery itself is an adaptive process. They do not expect people to move seamlessly between changing conditions simply because the equipment has been repaired. They make room for learning on the way back just as they make room for learning on the way through.

Perhaps that is the final lesson this ordinary week had to offer. We can reduce uncertainty. We can prepare for it. We can build systems that respond to it with remarkable grace. But we cannot eliminate it. Every solution reveals another question, every recovery begins another adjustment, and every return to normal becomes the starting point for a new reality.

Uncertainty is not the enemy of operational maturity.

It is the environment in which operational maturity proves itself.

Graphic Elements and Callouts