CSF Building a Metrics Program – Recover Part 6 of 6
Recover isn’t about disaster recovery, its about restoring business to an acceptable operating state and if we learned enough to become resilient.
You’re on the final part, Recover, of my series on Building a Metrics Program using the CSF. I wrote about Govern, Identify, Protect, Detect, and Respond prior to this page, leveraging the CSF to identify what is important to manage an operational team and to report to the board or any executive committee.

Using NIST CSF 2.0, Recover has three categories:
- RC.RP — Incident Recovery Plan Execution
- RC.CO — Incident Recovery Communication
- RC.IM — Incident Recovery Improvements
While there may be several metrics, having a regular review of the incident response plan gets everyone who should be involved, at least familiar with what’s going on. Reviewing steps with people who are called out in the plan, keeps it in their mind that they are accountable or responsible for the success of their own programs and that cybersecurity isn’t a bullet proof vest that can be used for every attack that comes through. Business still has to run and it should run despite hiccups or disasters in the business’ underlying layer.
Recovery’s Incident Recovery Plan Execution confirms that restoration of business services within the time and level of availability meets what the business requires.
| Control / Outcome | Leading indicators | Lagging indicators |
|---|---|---|
| RC.RP-01 Recovery plan is executed during recovery | % critical services with current recovery plans; recovery exercise completion; plan activation readiness | Recovery plans that fail during actual incidents |
| RC.RP-02 Recovery actions are selected and prioritized according to recovery criteria | % critical services with defined RTO/RPO and recovery priorities; recovery decisions mapped to business requirements | Recovery priorities inconsistent with business needs; critical services restored in the wrong order |
| RC.RP-03 Backup systems and assets are verified before use | % critical systems with tested backups; backup integrity verification rate; immutable/offline backup coverage | Failed restoration; corrupted/unusable backups |
| RC.RP-04 Critical systems are restored | % critical services meeting recovery objectives during exercises; recovery readiness coverage | Actual recovery time vs RTO; services failing to meet recovery objectives |
| RC.RP-05 Recovery operations are verified before returning to normal operations | % recovered systems undergoing security validation; recovery validation completion rate | Reintroduced vulnerabilities; compromised systems returning to production |
| RC.RP-06 Recovery activities are completed according to defined criteria | % recovery activities meeting established criteria; business-owner validation rate | Services returned to operation before being adequately secured or validated |
While the jobs are split between IT and Business, its up to Cybersecurity to help measure and write up the gaps and help both partners with creating a place for them to plan how to close the gaps. Something to report is what % of critical business services can be restored within the business’ defined recovery objectives.
Recovery’s Incident Recovery Communication informs leaders, employees, customers, partners, and other stakeholders where you are in the recovery process. It guides them along as some incidents are unavoidable, this defines your resiliency.
| Control / Outcome | Leading indicators | Lagging indicators |
|---|---|---|
| RC.CO-01 Recovery activities and progress are communicated to designated internal and external stakeholders | % material recoveries with established communication cadence; recovery status reporting compliance | Stakeholders receiving delayed/inaccurate recovery information |
| RC.CO-02 Public communications are managed according to established criteria | % applicable scenarios with preapproved communication plans; spokesperson readiness; exercise validation | Inconsistent or delayed public communication during actual events |
| RC.CO-03 Recovery activities are coordinated with internal and external parties | % critical recovery dependencies with identified contacts; supplier recovery exercises | Recovery delays caused by coordination failures |
| RC.CO-04 Recovery status is communicated to leadership and decision-makers | % recovery milestones reported within defined timeframe; executive reporting compliance | Leadership making decisions using outdated recovery information |
The metric isn’t about how well you communicate, but how many decisions are required in the recovery process to quickly answer service status, expected restoration time, business impact, and handling remaining risk. It’s decision visibility.
Recovery’s Incident Recovery Improvements connects back to Identify, Protect, and Respond… Did we become more resilient because of what happened? Did we find root cause and fix it so it doesn’t happen again? If nothing changes, you’ve potentially paid the cost of an incident without getting the benefit of the lesson.
| Control / Outcome | Leading indicators | Lagging indicators |
|---|---|---|
| RC.IM-01 Recovery plans incorporate lessons learned | % recovery plans updated after incidents/exercises; % lessons with assigned owners and target dates | Repeat recovery failures; recovery plans failing in subsequent incidents |
| RC.IM-02 Recovery strategies are updated based on lessons learned | % material recovery findings resulting in strategy changes; corrective-action completion rate | Recurring business disruption; repeated recovery weaknesses |
If this is a recovery from a type of attack that happens more than once, metrics of how long the incident took can be measured on it’s improvement.
For the executive dashboard, we should be able to answer the following in normal non-technobabble people-speak.
| Metric | Executive question |
|---|---|
| 1. Critical Business Service Recovery Achievement | Are critical services restored within their required timeframe? |
| 2. RTO Achievement | Are we recovering quickly enough? |
| 3. RPO Achievement | Are we recovering enough data? |
| 4. Backup Recoverability Coverage | Can we actually restore from our backups? |
| 5. Recovery Validation Coverage | Have recovered systems been verified as secure and operational? |
| 6. Recovery Dependency Coverage | Do we understand everything required to restore critical services? |
| 7. Recovery Decision Visibility | Does leadership have the information required to make recovery decisions? |
| 8. Third-Party Recovery Readiness | Can critical suppliers recover when we need them? |
| 9. Recovery Improvement Closure Rate | Are lessons learned actually being implemented? |
| 10. Resilience Improvement Effectiveness | Are those improvements measurably making us more resilient? |
