CSF Building a Metrics Program – Respond Part 5 of 6

Another program near to me as I took over the SOC. The NIST and the CSF guided me in the ways of incident response… that and my kick-ass team who built the program with me. Without a doubt, since you are on part 5 of 6, I feel you are prepared to read more, but in case you need to jump back, feel free to read about Govern, Identify, Protect, and Detect, of my series, Building a Metrics Program.

Respond is about helping the organization move from once we know something bad happened to controlling the situation and reducing any risk impact… well, keeping it contained if possible.

Using NIST CSF 2.0, Respond has five categories:

  • RS.MA — Incident Management
  • RS.AN — Incident Analysis
  • RS.CO — Incident Response Reporting and Communication
  • RS.MI — Incident Mitigation
  • RS.IM — Incident Response Improvement

When a negative security event occurs, can we make good decisions quickly to limit its risk impact? Measuring these metrics will help answer that.

Respond’s Incident Management organizes and manages a cybersecurity incident quickly, establishes ownership, and coordinates the people needed to make decisions on actions that need to be taken.

Control / OutcomeLeading indicatorsLagging indicators
RS.MA-01 The incident response plan is executed with relevant third parties once an incident is declared% material scenarios with tested response plans; exercise completion; third-party participationIncidents where response plans were unavailable, outdated or ineffective
RS.MA-02 Incident reports are triaged and validated% incidents triaged within SLA; triage accuracy; escalation accuracyMaterial incidents initially misclassified or escalated late
RS.MA-03 Incidents are categorized and prioritized% incidents categorized according to defined criteria; prioritization QAMaterial incidents assigned incorrect priority; delayed response to high-impact incidents
RS.MA-04 Incidents are escalated according to defined criteria% incidents meeting escalation criteria that are escalated within SLA; escalation-path validationLate executive/business notification; missed escalation
RS.MA-05 Criteria for initiating incident response are established and enforced% response scenarios with defined activation criteria; exercise validationDelayed incident declaration; incidents managed informally before formal activation

Another thing to measure, if you can, is when the alert is received to when the first action was performed.

Respond’s Incident Analysis when what happened was clearly understood, criticality was confirmed, impact was figured out, and next steps are determined.

Control / OutcomeLeading indicatorsLagging indicators
RS.AN-01 Investigations establish what happened% material incidents with documented investigation; investigation SLA adherenceUnknown root cause; incomplete incident understanding
RS.AN-02 Actions and events are recorded with integrity% material incidents with complete timeline; evidence preservation complianceInability to reconstruct incident timeline; lost/insufficient evidence
RS.AN-03 Analysis is performed to establish what has been affected% material incidents with documented scope assessment; critical asset/business-service context availableScope discovered late; affected systems/services underestimated
RS.AN-04 Incidents are categorized according to defined criteriaClassification accuracy; incident taxonomy coverageIncidents reclassified after material impact becomes apparent
RS.AN-05 Forensics are performed when appropriate% applicable incidents receiving forensic analysis; evidence preservation complianceRoot cause remains unknown; forensic evidence unavailable
RS.AN-06 Information is provided to authorized personnel and relevant stakeholders% material incidents with current situation reports; stakeholder notification SLABusiness leaders receive incomplete or delayed information
RS.AN-07 Incident estimates and forecasts are updated as new information becomes available% material incidents with updated impact estimates; forecast update frequencySignificant variance between initial and final impact estimates

Calculate the time required to establish the affected assets, accounts, data, business services, and third parties associated with the incident. That answers “How big is this!?”

Respond’s Incident Response Reporting & Communication confirms if the right people are getting the right information at the right time so decisions are timely.

Control / OutcomeLeading indicatorsLagging indicators
RS.CO-01 Personnel know their roles and responsibilities during incidents% response roles trained; exercise participation; role acknowledgmentConfusion over ownership during incidents
RS.CO-02 Internal and external stakeholders are notified according to response plans% required notifications completed within SLA; notification accuracyLate notification; missed notification requirements
RS.CO-03 Information is shared according to incident response plans% material incidents with established communication cadence; situation-report complianceStakeholders operating from conflicting information
RS.CO-04 Coordination with designated internal and external stakeholders occurs% material scenarios with identified contacts; exercise participationDelayed coordination with legal, privacy, regulators, law enforcement, vendors or business leadership
RS.CO-05 Incident information is communicated to leadership and decision-makers% material incidents with executive updates within defined intervals; decision log completenessExecutive decisions delayed by inadequate information

This isn’t always a metric until the post-mortem if the result was catastrophic, so hand down to your teams how important it is to measure how much data did the incident responder provide for the responsible executive to create an informed decision. It might have to be preceded with inclusion of the right people when creating or testing playbooks.

e.g. When I was in healthcare, before I could make a determination or recommend a determination to cut a B2B connection with a business associate, I needed to know if the connection was providing life saving data or administrative data. Doing the right thing from a cybersecurity perspective might not be doing the right thing from a people endangerment perspective.

Respond’s Incident Mitigation determines that during an incident, we are minimizing potential damage.

Control / OutcomeLeading indicatorsLagging indicators
RS.MI-01 Incidents are contained% material incidents with containment strategy initiated within target; containment SLA adherenceTime from incident declaration to containment; spread of compromise before containment
RS.MI-02 Incidents are eradicated% material incidents with eradication plan; eradication completion rateRecurrence of compromised systems/accounts; reinfection
RS.MI-03 Newly identified vulnerabilities are mitigated or documented% incident-related vulnerabilities addressed; emergency remediation completionRe-exploitation of known weaknesses; incidents recurring through same vulnerability
RS.MI-04 Incident mitigation activities are validated% containment/eradication actions validated; post-action testing coverageFailed containment; attacker persistence after declared containment

Measure when the incident is announced to when everyone can stand down from the call. Containment has happened and is validated and those responsible for recovery have tickets or duties assigned to them to happen as soon as possible or within a specified time.

Respond’s Incident Response Improvement asks the tough and necessary post-mortem questions.

Control / OutcomeLeading indicatorsLagging indicators
RS.IM-01 Response plans incorporate lessons learned% material incidents producing lessons learned; % response plans updatedRepeat response failures; same response weakness recurring
RS.IM-02 Response strategies are updated based on lessons learned% material findings with assigned corrective actions; corrective-action completion rateRecurring incidents with same root cause; repeat control failures

Did the way we work to respond work in our favor? Could anything be improved? Did we delay resolution in any way?

For an executive dashboard, I’d put something like this together.

MetricExecutive question
1. Incident Activation TimeHow quickly do we formally mobilize?
2. Incident Scope Determination TimeHow quickly do we understand the size of the problem?
3. Business Impact Assessment TimeHow quickly do we know what the event means to the business?
4. Decision-Ready Information RateAre decision-makers getting what they need?
5. Executive Escalation TimelinessAre material events reaching leadership quickly enough?
6. Business Impact Containment TimeHow quickly do we stop the damage from growing?
7. Containment EffectivenessDid containment actually stop attacker activity/spread?
8. Eradication EffectivenessDid we actually remove the threat?
9. Incident Impact vs. EstimateHow accurate were our initial assumptions?
10. Response Recurrence RateAre we actually getting better?

Related Posts