🇯🇵 日本語 🇬🇧 English 🇨🇳 中文 🇲🇾 Bahasa Melayu

Disaster-Revealed IT Exposes Management’s Lack of “Production-Ready” Thinking

What Systems That Work Even Without Connectivity Are Asking

A news story about “a digital system that keeps running even when communications go down—how it could transform evacuation center reception” has been generating buzz. It details a proof-of-concept experiment for a system that digitalizes evacuation center check-ins and operates offline.

At first glance, this seems like just another example of disaster-related digital transformation (DX). However, from a management × IT perspective, it encapsulates a fundamental problem many companies face.

That problem is the lack of “IT that doesn’t stop in a production environment.” Many companies implement systems only thinking about daily operational efficiency. But what truly holds value is an IT infrastructure that continues to function even when trouble strikes.

The Danger of Separating Normal and Disaster Scenarios

The key feature of this evacuation center system is that it operates standalone even if communications are cut off. It doesn’t rely on the cloud; it completes check-ins using a local database. This holds extremely important implications for management as well.

Many companies adopt a two-pronged IT strategy: “cloud for normal times, backup for disasters.” But this very mindset is dangerous. Why? Because disasters strike suddenly, and smoothly switching systems isn’t always possible.

In fact, an IT manager at a mid-sized manufacturing company shared this: “Even though we have a BCP (Business Continuity Plan), the recovery procedures for when the system actually goes down were only documented on paper. I have no confidence we could get it running when it really matters.”

This is a classic example of treating IT only as “administrative IT.” If management had designed systems directly linked to sales and customer service as “business IT” that wouldn’t stop in a production environment, this problem could have been prevented.

The “Offline-First” Design Philosophy Management Should Consider

What the evacuation center case shows is that we need to change the very philosophy of system design.

Specifically, management should focus on the following three points.

1. Don’t Evaluate System Availability Separately for “Normal Times” and “Failure Times”

Many companies evaluate system uptime as “99.9%.” However, it’s the remaining 0.1% of failure time that can have a fatal impact on customer service and operations.

Management should demand “full-spec availability” from systems, including during failures. Specifically, consider adding an “offline-first” design philosophy—where core functions work even offline—as a condition when selecting SaaS.

2. Design Backup as a “Goal,” Not Just a “Means”

Many companies think of data backup as merely “taking it just in case for peace of mind.” But what’s truly necessary is “being able to recover from the backup,” and the recovery procedure must not be dependent on specific individuals.

For example, at a major retail chain, when the POS system went down, cashiers would write out slips by hand and manually enter them into the system later. This is a classic case of not designing the “operations” for when the system stops. Management needs to evaluate the effectiveness of IT investments including the “business processes” for when the system is down.

3. Ensure “Reproducibility” as Business IT

Not just during disasters, but also for everyday troubles (server crashes, network failures, cyberattacks), “reproducibility” is essential to keep operations running.

Reproducibility means a state where anyone can continue operations using the same procedure, without relying on a specific person. This is the core of business IT. Just as the digitalization of evacuation center reception moved away from “paper ledgers” and allowed anyone to complete check-ins in the same way, a company’s core business operations should also aim for a state where “anyone can achieve the same result” during a failure.

Concrete Example: A Logistics Company’s “Production-Ready” Failure

There was a case where a logistics company revamped its warehouse management system (WMS). They introduced a cloud-based WMS, achieving cost reduction and real-time inventory management. However, six months after implementation, a data center failure caused the system to go down for 12 hours.

During that time, all picking operations in the warehouse stopped. Shipments were delayed, and the company faced significant damage claims from its business partners. The management of this company had completely failed to consider “alternative measures during failures” when introducing the system.

This case teaches us that “business continuity during failures” must always be included as a criterion for IT investment decisions. Beyond just the features and price of a tool, evaluation criteria should include “Does it work offline?” and “Are alternative processes prepared for failures?”

3 Actions Management Should Take Right Now

So, what specifically should management do? Here are three suggestions.

1. Review Your Core System’s “Failure Operations Manual”
Don’t leave it entirely to the IT department. Management itself should quantitatively assess “how much impact a system outage would have on sales and customer service.” Based on that, define the recovery time objective (RTO) and the alternative business processes to be implemented during that time.

2. Add “Offline Functionality” to Your SaaS Selection Criteria
When introducing new SaaS, always ask the vendor “which features will work if communications are cut off.” Salespeople will tout “99.9% availability,” but it’s management’s role to ask what happens during that remaining 0.1%.

3. Strengthen Reproducibility as “Business IT”
Document and standardize system settings and operational rules that only specific people know. This is effective not only during disasters but also to prevent operational stagnation due to employee resignations or transfers. Specifically, introduce knowledge management tools like Notion or Confluence to create a state where anyone can access the information.

Conclusion: IT’s Greatest Value Is “Not Stopping”

What the digitalization of evacuation center reception shows is an obvious yet often overlooked fact: before being “convenient,” IT must be “uninterruptible.”

When management judges IT investments, they tend to focus on “how much will operational efficiency improve?” or “how much can costs be reduced?” However, the real question to ask is “What will happen to the company if this system stops?”

Any IT investment that can’t answer this question is merely “decorative IT” that won’t function in a production environment. To continue providing value to customers during disasters and troubles, management must directly design “production-ready” scenarios and embed them into IT.

Is your company’s IT truly designed to “not stop”?

Comments

Copied title and URL