CSF Building a Metrics Program – Recover Part 6 of 6

Recover isn’t about disaster recovery, its about restoring business to an acceptable operating state and if we learned enough to become resilient.

You’re on the final part, Recover, of my series on Building a Metrics Program using the CSF. I wrote about Govern, Identify, Protect, Detect, and Respond prior to this page, leveraging the CSF to identify what is important to manage an operational team and to report to the board or any executive committee.

Using NIST CSF 2.0, Recover has three categories:

  • RC.RP — Incident Recovery Plan Execution
  • RC.CO — Incident Recovery Communication
  • RC.IM — Incident Recovery Improvements

While there may be several metrics, having a regular review of the incident response plan gets everyone who should be involved, at least familiar with what’s going on. Reviewing steps with people who are called out in the plan, keeps it in their mind that they are accountable or responsible for the success of their own programs and that cybersecurity isn’t a bullet proof vest that can be used for every attack that comes through. Business still has to run and it should run despite hiccups or disasters in the business’ underlying layer.

Recovery’s Incident Recovery Plan Execution confirms that restoration of business services within the time and level of availability meets what the business requires.

Control / OutcomeLeading indicatorsLagging indicators
RC.RP-01 Recovery plan is executed during recovery% critical services with current recovery plans; recovery exercise completion; plan activation readinessRecovery plans that fail during actual incidents
RC.RP-02 Recovery actions are selected and prioritized according to recovery criteria% critical services with defined RTO/RPO and recovery priorities; recovery decisions mapped to business requirementsRecovery priorities inconsistent with business needs; critical services restored in the wrong order
RC.RP-03 Backup systems and assets are verified before use% critical systems with tested backups; backup integrity verification rate; immutable/offline backup coverageFailed restoration; corrupted/unusable backups
RC.RP-04 Critical systems are restored% critical services meeting recovery objectives during exercises; recovery readiness coverageActual recovery time vs RTO; services failing to meet recovery objectives
RC.RP-05 Recovery operations are verified before returning to normal operations% recovered systems undergoing security validation; recovery validation completion rateReintroduced vulnerabilities; compromised systems returning to production
RC.RP-06 Recovery activities are completed according to defined criteria% recovery activities meeting established criteria; business-owner validation rateServices returned to operation before being adequately secured or validated

While the jobs are split between IT and Business, its up to Cybersecurity to help measure and write up the gaps and help both partners with creating a place for them to plan how to close the gaps. Something to report is what % of critical business services can be restored within the business’ defined recovery objectives.

Recovery’s Incident Recovery Communication informs leaders, employees, customers, partners, and other stakeholders where you are in the recovery process. It guides them along as some incidents are unavoidable, this defines your resiliency.

Control / OutcomeLeading indicatorsLagging indicators
RC.CO-01 Recovery activities and progress are communicated to designated internal and external stakeholders% material recoveries with established communication cadence; recovery status reporting complianceStakeholders receiving delayed/inaccurate recovery information
RC.CO-02 Public communications are managed according to established criteria% applicable scenarios with preapproved communication plans; spokesperson readiness; exercise validationInconsistent or delayed public communication during actual events
RC.CO-03 Recovery activities are coordinated with internal and external parties% critical recovery dependencies with identified contacts; supplier recovery exercisesRecovery delays caused by coordination failures
RC.CO-04 Recovery status is communicated to leadership and decision-makers% recovery milestones reported within defined timeframe; executive reporting complianceLeadership making decisions using outdated recovery information

The metric isn’t about how well you communicate, but how many decisions are required in the recovery process to quickly answer service status, expected restoration time, business impact, and handling remaining risk. It’s decision visibility.

Recovery’s Incident Recovery Improvements connects back to Identify, Protect, and Respond… Did we become more resilient because of what happened? Did we find root cause and fix it so it doesn’t happen again? If nothing changes, you’ve potentially paid the cost of an incident without getting the benefit of the lesson.

Control / OutcomeLeading indicatorsLagging indicators
RC.IM-01 Recovery plans incorporate lessons learned% recovery plans updated after incidents/exercises; % lessons with assigned owners and target datesRepeat recovery failures; recovery plans failing in subsequent incidents
RC.IM-02 Recovery strategies are updated based on lessons learned% material recovery findings resulting in strategy changes; corrective-action completion rateRecurring business disruption; repeated recovery weaknesses

If this is a recovery from a type of attack that happens more than once, metrics of how long the incident took can be measured on it’s improvement.

For the executive dashboard, we should be able to answer the following in normal non-technobabble people-speak.

MetricExecutive question
1. Critical Business Service Recovery AchievementAre critical services restored within their required timeframe?
2. RTO AchievementAre we recovering quickly enough?
3. RPO AchievementAre we recovering enough data?
4. Backup Recoverability CoverageCan we actually restore from our backups?
5. Recovery Validation CoverageHave recovered systems been verified as secure and operational?
6. Recovery Dependency CoverageDo we understand everything required to restore critical services?
7. Recovery Decision VisibilityDoes leadership have the information required to make recovery decisions?
8. Third-Party Recovery ReadinessCan critical suppliers recover when we need them?
9. Recovery Improvement Closure RateAre lessons learned actually being implemented?
10. Resilience Improvement EffectivenessAre those improvements measurably making us more resilient?

Related Posts