Your risk register says the risk is under control. Your recovery plan promises two hours. But when the system goes down, can you prove either claim?
Image: AI GeneratedThe Incident That Exposed a Bigger Problem
It was an ordinary working day at a bank.
Customers were making payments, branches were processing transactions, and digital banking services were operating as expected. Behind the scenes, the Information Security team had completed its risk assessment. The risk register had been reviewed, critical systems had been identified, and the disaster recovery plan had received management approval.
Everything appeared to be under control.
Then the core banking system became unavailable.
The technical team began investigating. Initial checks did not reveal an obvious cause. Transaction processing slowed, branches struggled to serve customers, and complaints started arriving through multiple channels.
The incident continued.
After two hours, the recovery target stated in the disaster recovery plan had been missed. At four hours, the business impact had become unacceptable according to the bank's business continuity requirements. Yet the system remained unavailable.
Six hours after the incident began, the service was finally restored.
An uncomfortable question emerged during the post-incident review:
How could an organization with an approved risk register, a documented Business Impact Analysis, and a disaster recovery plan still fail to recover within its required timeframe?
The answer might not be a missing security control or an inadequate recovery technology. It might be a fundamental disconnect between risk assessment, business requirements, and what the organization had actually demonstrated it could do.
This fictional scenario illustrates a problem that security practitioners should take seriously: documenting a risk is not the same as understanding it, and defining a recovery target is not the same as proving that it can be achieved.
1. The Risk Register: Evidence or Professional Guesswork?
A risk register is one of the most familiar documents in information security management. It records identified risks, their likelihood and impact, existing controls, risk owners, treatment decisions, and residual risk.
However, the existence of a risk register does not automatically mean that an organization understands its risks.
Consider a common entry:
Risk: Core banking system outage.
Likelihood: Medium.
Impact: High.
Overall risk: High.
Existing controls: Backup, disaster recovery site, and monitoring.
At first glance, this looks reasonable. But ask a few additional questions:
- What evidence supports the likelihood rating?
- How effective are the existing controls?
- Has the disaster recovery site been tested under realistic failure conditions?
- Can the organization restore the service within its required recovery time?
- Does the impact rating reflect the consequences of an outage after 30 minutes, two hours, four hours, or an entire business day?
If the risk owner cannot answer these questions, the assessment may be based partly on assumptions rather than demonstrable evidence.
Professional judgment is unavoidable in risk management. Historical data may be limited, emerging threats may have no reliable frequency estimates, and the effectiveness of a control may be uncertain. The problem is not judgment itself. The problem is undocumented judgment presented as established fact.
A defensible risk assessment should identify the evidence used, explain the rationale for the rating, document assumptions, and acknowledge uncertainty.
For a core banking system, useful evidence could include historical outage records, infrastructure architecture, single points of failure, vulnerability findings, control testing, disaster recovery exercise results, and third-party service dependencies.
This evidence does not eliminate uncertainty. It makes the assessment more transparent, repeatable, and open to challenge.
The difference between a rating and a reason
A red risk rating communicates a conclusion. It does not explain how that conclusion was reached.
Likewise, a green control status does not prove that a control is effective under the conditions that matter.
A backup process might report successful completion every night, yet the organization might never have tested whether the backups could restore the complete service within the required timeframe. A disaster recovery site might exist, yet dependencies on identity services, network connectivity, encryption keys, or external payment providers could prevent successful recovery.
The security practitioner must therefore look beyond the status recorded in the register and ask whether the supporting evidence justifies it.
But risk assessment alone does not tell us exactly how much disruption the business can tolerate. For that, we need to understand the business itself.
2. Business Impact Analysis: Understanding What Failure Means
Business Impact Analysis (BIA) changes the perspective from technical failure to business consequences.
A technical team might report that a database server is unavailable. The business, however, experiences delayed transactions, unavailable services, operational backlogs, customer complaints, potential financial losses, and possibly regulatory or contractual consequences.
The technical incident and the business impact are connected, but they are not the same thing.
A BIA identifies critical business activities, assesses the consequences of their disruption over time, and examines the resources and dependencies required to resume them.
For a bank, this may include:
- Core banking and transaction processing.
- Digital banking and customer authentication.
- Payment clearing and settlement.
- ATM and card services.
- Customer support and branch operations.
- Supporting identity, network, database, and third-party services.
The impact of a disruption may change significantly as time passes.
An outage lasting 15 minutes might cause a manageable number of failed transactions. An outage lasting several hours could create substantial backlogs, affect settlement deadlines, and prevent customers from accessing essential services. A prolonged disruption might introduce additional financial, contractual, regulatory, and reputational consequences.
The BIA helps the organization understand these changing impacts and determine which activities require the highest recovery priority.
It also identifies dependencies that are easily overlooked when assessments focus only on individual systems.
For example, restoring a banking application is not enough if its database is unavailable, its authentication service cannot be reached, or its network dependencies have not been recovered.
A system's technical importance and a business activity's criticality are related, but they are not interchangeable.
The BIA provides the business context needed to establish recovery requirements.
3. MTPD: How Long Can the Business Survive the Disruption?
One of the most important outcomes of a BIA is the determination of the Maximum Tolerable Period of Disruption (MTPD).
MTPD represents the maximum duration of a disruption before the resulting impacts become unacceptable to the organization.
It is not simply the amount of downtime that management would prefer to avoid. It is a business tolerance that should be established through impact analysis, documented requirements, and appropriate approval.
Return to our fictional banking incident.
Suppose the BIA determines that a particular critical business activity cannot remain disrupted for more than four hours without unacceptable consequences.
The activity's MTPD is therefore four hours.
This does not mean the organization should plan to restore the service at the four-hour mark. MTPD is the outer limit of tolerable disruption, not the operational recovery target.
The distinction matters because recovery is rarely instantaneous. Even after a service becomes available, the organization may need to validate transactions, reconcile records, clear backlogs, and confirm that business operations can resume safely.
The MTPD should therefore guide the recovery strategy rather than become an excuse to delay recovery until the last possible moment.
4. RTO and RPO: Turning Business Needs into Recovery Objectives
Once the organization understands the consequences of disruption, it can establish appropriate recovery objectives.
Two terms are especially important: Recovery Time Objective (RTO) and Recovery Point Objective (RPO).
RTO defines the target time for restoring a business activity or service following a disruption.
RPO defines the point in time to which data must be recovered, expressed as a period of time representing the targeted maximum amount of data loss.
Consider the following illustrative requirements for the banking service:
| Measure | Example | Meaning |
|---|---|---|
| MTPD | 4 hours | Disruption must not exceed this limit before impacts become unacceptable |
| RTO | 2 hours | The target is to restore the service within two hours |
| RPO | 15 minutes | Recovery must meet the agreed data-loss tolerance of 15 minutes |
These values are illustrative, not universal banking requirements.
The two-hour RTO gives the organization a target that is shorter than the four-hour MTPD. The 15-minute RPO addresses a different concern: how much data loss can be tolerated following a disruptive event.
Meeting one objective does not guarantee meeting the others.
A service might be restored within two hours but lose more data than the RPO permits. Alternatively, data might be recovered to an acceptable point, but the service could remain unavailable for six hours.
Both scenarios represent recovery failures against the specified objectives.
The organization must therefore validate its recovery strategy against both business and technical requirements.
5. The Uncomfortable Truth: Your RTO Is Not Evidence
Now return to the incident.
The bank's recovery plan specifies an RTO of two hours. However, the last realistic recovery exercise took six hours.
Which number should the security practitioner trust?
The two-hour RTO describes what the organization intends to achieve. The six-hour test describes what happened under the conditions of that exercise.
The difference is a material capability gap that requires investigation.
Perhaps the recovery environment was not adequately provisioned. Perhaps restoring the database took longer than expected. Perhaps critical dependencies were missing from the recovery sequence. Perhaps the recovery team lacked the access permissions or documented procedures needed to complete the work efficiently.
The test result does not automatically mean that every future recovery will take six hours. Different scenarios can produce different outcomes. Nevertheless, it is evidence that the organization has not yet demonstrated that it can reliably achieve the two-hour objective under the tested conditions.
The gap should be documented, assigned to an accountable owner, assessed for risk, and addressed through corrective action.
Potential measures include improving automation, removing single points of failure, validating backup restoration, correcting dependency issues, improving recovery procedures, and conducting further exercises.
The important point is that recovery objectives must be supported by credible technical evidence.
A recovery plan describes the intended capability. A recovery test provides evidence about actual capability.
Neither should be confused with the other.
6. Connecting the Risk Register, BIA, and Business Continuity
The relationship between these activities is often misunderstood.
They are not simply three documents to complete for an audit. They are complementary parts of a broader management process.
The risk register identifies and evaluates risks, records existing controls, and tracks treatment decisions.
The BIA assesses business consequences, identifies critical activities and dependencies, and establishes recovery requirements.
MTPD defines the maximum tolerable period of disruption. RTO and RPO translate recovery needs into specific time-based objectives.
Recovery exercises then test whether the proposed arrangements can meet those objectives.
The results feed back into risk assessment and continuity improvement.
Consider the relationship in practice:
- Identify the risk. A failure or cyberattack could make a critical banking service unavailable.
- Assess the business impact. The BIA determines how the consequences develop as disruption continues.
- Establish recovery requirements. The business approves an MTPD and appropriate RTO and RPO.
- Evaluate capability. Technical teams assess whether the architecture, backups, dependencies, and recovery procedures can meet those requirements.
- Record and treat gaps. Where recovery capability is insufficient, the risk register records the issue, treatment plan, owner, and target date.
- Validate improvements. Recovery exercises provide evidence of whether corrective actions have improved the organization's capability.
This is not necessarily a linear process. Risk assessments can reveal business dependencies that require a new BIA analysis, while BIA findings can reveal risks that were not previously recognized. The activities should inform one another and be revisited when systems, threats, business processes, or dependencies change.
The aim is not to force every document to contain identical information. It is to ensure that their conclusions are consistent and traceable.
7. What Security Practitioners Should Challenge
When reviewing a risk register or continuity plan, practitioners should move beyond asking whether the required document exists.
Ask whether the organization can demonstrate the reasoning behind its decisions.
When reviewing the risk register:
- Is the likelihood rating supported by evidence or documented assumptions?
- Are existing controls tested for effectiveness?
- Is residual risk assessed after considering those controls?
- Are treatment decisions assigned to accountable owners?
When reviewing the BIA:
- Have the relevant business owners validated the impact assessment?
- Are financial, operational, regulatory, contractual, and reputational impacts considered?
- Have critical dependencies and third-party services been identified?
- Are the recovery requirements consistent with business needs?
When reviewing MTPD, RTO, and RPO:
- Is the MTPD supported by the impact analysis?
- Is the RTO shorter than the MTPD, with a realistic margin for recovery and stabilization?
- Does the RPO reflect the business's tolerance for data loss?
- Have the objectives been translated into achievable technical requirements?
When reviewing recovery capability:
- Have backups actually been restored and validated?
- Have recovery exercises tested realistic failure scenarios?
- Are identity, network, database, and external service dependencies included?
- Do test results demonstrate that recovery objectives can be achieved?
- Are identified gaps tracked through to closure?
These questions help shift security assurance from document completion to evidence-based evaluation.
8. The Role of ISO 27001 and ISO 22301
This approach also reflects the complementary purposes of two widely used management system standards.
ISO/IEC 27001 provides a framework for establishing and continually improving an Information Security Management System (ISMS), including information security risk assessment and treatment.
ISO 22301 provides a framework for establishing and continually improving a Business Continuity Management System (BCMS), including understanding disruption impacts and developing continuity and recovery capabilities.
The standards address different but related management needs. Information security risk management helps organizations understand and treat information security risks. Business continuity management focuses on maintaining or recovering prioritized activities during disruption.
A ransomware incident, for example, may create information security risks involving confidentiality, integrity, and availability. At the same time, the resulting service outage may threaten the organization's ability to continue critical business activities within their tolerable disruption periods.
Integrating the two perspectives helps the organization assess not only whether a cyberattack is possible, but also whether it can sustain or recover critical operations when one occurs.
However, simply maintaining a risk register, a BIA, and a recovery plan does not establish conformity with either standard. Organizations must implement the relevant requirements, operate their processes, and evaluate their effectiveness within the applicable scope.
9. From Compliance to Demonstrable Resilience
The banking scenario began with an organization that appeared prepared. It had documented its risks, approved its recovery objectives, and maintained a disaster recovery plan.
What it lacked was sufficient evidence that its recovery capability matched its business requirements.
This is where security practitioners can make a meaningful difference.
Challenge unsupported risk ratings. Ask business owners to validate impact assessments. Verify that recovery objectives are achievable. Review the results of realistic exercises. Ensure that gaps are recorded, assigned, and resolved.
Where uncertainty remains, document it honestly and make it part of the decision-making process.
Not every risk can be eliminated, and not every recovery scenario can be predicted. The objective is not to create a perfect risk register. It is to build a defensible understanding of uncertainty, business impact, and recovery capability.
A risk register should not be a collection of colored cells. A BIA should not be a spreadsheet completed merely to satisfy an audit. A recovery plan should not be a promise unsupported by testing.
Together, these practices should help answer four essential questions:
- What could go wrong?
- How would the business be affected?
- How long can the disruption be tolerated?
- Can the organization recover within the required limits?
If those answers are based on evidence, connected through accountable decisions, and validated through testing, the organization is moving beyond paperwork towards resilience.
Because resilience is not what your documents promise. It is what your organization can demonstrate.
References
- International Organization for Standardization. (2018). ISO 31000:2018 risk management—Guidelines. https://www.iso.org/standard/65694.html
- International Organization for Standardization. (2019). ISO 22301:2019 security and resilience—Business continuity management systems—Requirements. https://www.iso.org/standard/75106.html
- International Organization for Standardization. (2022). ISO/IEC 27001:2022 information security, cybersecurity and privacy protection—Information security management systems—Requirements. https://www.iso.org/standard/27001
- Ross, R. (2012). Guide for conducting risk assessments (NIST Special Publication 800-30 Rev. 1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-30r1
- Swanson, M., Bowen, P., Phillips, A. W., Gallup, D., & Lynes, D. (2010). Contingency planning guide for federal information systems (NIST Special Publication 800-34 Rev. 1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-34r1
