It covered vendor failures with alternative supplier provisions. What it did not adequately cover was a sustained outage of the cloud provider whose infrastructure hosted eighty percent of the organization's production workloads. That scenario was mentioned in a single paragraph. The paragraph described the scenario as 'unlikely given the provider's redundancy architecture.' The outage lasted fourteen hours. The recovery took three days.
The Dependency That BCP Didn't Model
Cloud adoption has fundamentally changed the business continuity risk landscape in ways that many BCP frameworks have not kept pace with. The pre-cloud BCP assumption — that the organization owns and operates the infrastructure it depends on, and therefore controls its availability — no longer holds for organizations whose primary infrastructure is cloud-hosted. The cloud provider's availability is the organization's availability. The cloud provider's recovery capability is the organization's recovery capability. Neither is fully within the organization's control.
This dependency is understood in general terms. What is less well understood in most BCP frameworks is its operational specificity: which workloads depend on which cloud services, what the recovery time objective for each workload is, how long the cloud provider's recovery would need to take before the organization's RTO is breached, and what the organization would actually do during a sustained provider outage. The BCP that addresses cloud provider outage in a single paragraph has not answered these questions.
Cloud vendor dependency is the most significant unmodeled risk in most business continuity programs. The provider's redundancy architecture does not prevent outages. It limits their frequency and scope. When the outage occurs, the BCP that relied on the provider's redundancy is the BCP that was not designed for the event that occurred.
What the Fourteen-Hour Outage Revealed
Workload Dependencies That Were Not Mapped
The incident response team's first task was understanding which systems were affected. The answer required mapping which production workloads ran on which cloud services — a mapping that did not exist in documented form. The cloud environment had grown organically, with workloads added continuously as engineering teams built new services. The dependency map was assembled during the outage, under operational pressure, with incomplete information. The fourteen-hour outage timeline included four hours of dependency mapping.
Failover That Was Tested in One Direction
The organization had a multi-region cloud architecture with primary and secondary regions. Failover procedures existed for regional failures. What was not tested was full provider-level failure — where both the primary and secondary regions of the same provider were affected simultaneously, as occurs in provider-wide networking or control plane outages. The tested failover scenario was regional. The actual outage was provider-wide.
Communication That Depended on the Failed Provider
The organization's incident communication infrastructure — the platform used to notify customers and internal stakeholders of the outage — was hosted on the same cloud provider experiencing the outage. During the first two hours of the incident, the communication platform was unavailable at the same time communication was most needed. The BCP included a communication procedure. It did not specify that the communication platform needed to be independent of the primary infrastructure provider.
Recovery Time That Exceeded the BCP's Assumption
The BCP's recovery time objective for primary application systems was four hours. The actual recovery time was three days. The gap between the RTO and the actual recovery time reflected two factors: the recovery took longer than expected because dependencies that were not fully mapped required discovery during recovery, and the cloud provider's own recovery took longer than the BCP's 'unlikely' scenario had assumed. The RTO was defined against a recovery scenario that was not the recovery scenario that occurred.
Building BCP for Cloud Dependency Reality
Business continuity planning for organizations with significant cloud dependency requires addressing three gaps that the traditional BCP model does not automatically cover.
Explicit cloud workload dependency mapping. A documented mapping of which production workloads depend on which cloud services, updated continuously as new workloads are deployed. This mapping is the prerequisite for understanding what is affected in a provider outage and for prioritizing recovery.
Provider-level outage as a modeled scenario. A specific BCP scenario that addresses sustained outage of the primary cloud provider, including the recovery procedures, the fallback options, the communication infrastructure that will function during the outage, and the tested recovery timeline for this specific scenario.
Multi-provider architecture for the highest-RTO workloads. For workloads whose RTO cannot tolerate the recovery time of a provider-level outage, a multi-provider architecture that enables continuity on a different provider while the primary recovers. This is not necessary for all workloads — it carries significant architectural complexity and cost. It is necessary for the workloads whose RTO makes provider-level dependency an unacceptable risk.
The cloud provider's redundancy reduces the probability of outage. It does not reduce the BCP requirement to model what happens when the outage occurs.
Model the cloud provider outage scenario explicitly. Test the recovery. Map the dependencies before the incident requires you to do it under pressure.
Enterprise practitioner perspective. Not legal advice. Part of the Deep Trust Governance Series by Verisq. Get the free weekly Breach Digest.
