Jump to section
You have prioritized your threat detection backlog, but now want to quantify the progress. Throughout my career leading multiple threat detection engineering teams, I have relied on well-defined metrics to measure outcomes, demonstrate program value, and drive continuous improvement. In this post, I outline the detection engineering metrics I consider most important, spanning detection quality and performance, detection maturity and engineering velocity, detection-as-code and CI/CD practices, threat hunting capacity, telemetry and data health, AI governance, modern SOC operations, incident triage efficiency, cloud identity posture, vulnerability remediation velocity, and Risk-Based Prioritization (RBP).

While some of these metrics are based on my own operational experience, many have been influenced by established security frameworks, industry best practices, and the work of other security leaders. I want to acknowledge and thank the contributors whose ideas and research have helped advance the detection engineering discipline. There are many more and as always, I will keep this post updated with anything new I learn. 🙂
List of Essential Detection Engineering Metrics
Risk-Based Prioritization (RBP) Metrics
This is a structured method for ranking risks by their likelihood and impact so organizations can focus resources on the most critical threats first. I like to use criteria such as EPSS/LEV for probability to determine which risks require immediate mitigation.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| EPSS/LEV Probability & Percentile | Exploit Prediction Scoring System/Likely Exploited Vulnerabilities model output (0.0 to 1.0 probability) | Predicts the probability of wild exploitation in the next 30 days. | Remediate immediately if EPSS > 0.20 or Percentile > 95% |
| CISA KEV Active Exploitation Flag | Binary flag (Yes/No) indicating presence in CISA KEV catalog | Identifies active, confirmed threat actor exploitation in the wild. | Remediate within 14 days (or mandatory federal SLA) |
| Prioritization Efficiency (Precision) | True exploited vulnerabilities prioritized/Total prioritized items | Measures how accurately the prioritization scheme targets real threats without wasted effort. | Maximize precision (avoid patching non-threats first) |
| Prioritization Coverage (Recall) | True exploited vulnerabilities prioritized/Total exploited vulnerabilities in environment | Ensures no actively exploited vulnerabilities are left out of the high-priority queue. | 100% recall on active wild exploits |
| Workload Reduction Rate | 1 – (High-Priority Items under RBP / High-Priority Items under CVSS >= 7.0) | Quantifies operational effort saved by moving from legacy CVSS to dynamic risk scoring. | 50% to 80% reduction in urgent remediation tasks |
| SSVC Decision Outcome Distribution | % allocation across SSVC buckets: Track, Track*, Attend, Act | Categorizes remediation actions based on vulnerability exploitation status and mission impact. | < 5% of total backlog in ‘Act’ bucket |
Vulnerability Remediation Velocity Metrics
This metric helps you identify the rate at which exploitable vulnerabilities are being reduced.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Mean Time to Patch (MTTP / MTTR) | Sum of remediated vulnerability durations / Total Closed Vulnerabilities | Traditional speed metric; heavily skewed by long-tail outliers (use with caution). | Tracked for regulatory compliance baseline |
| Median Time to Patch (p50) | 50th percentile duration of remediated vulnerabilities | Resistant to extreme outliers; represents the typical patch timeframe. | < 30 days for High/Critical vulnerabilities |
| High-Percentile SLA Metrics (p75, p90) | 75th and 90th percentile days to resolution | Exposes long-tail remediation delays in difficult legacy environments. | p90 < 60 days enterprise-wide |
| Vulnerability Half-Life | Days required to reduce an identified vulnerability cohort by 50% | Measures organizational remediation rate without assuming normal distribution. | < 21 days for critical asset cohorts |
| SLA Policy Compliance Rate | (Vulnerabilities remediated within policy window / total Identified) * 100 | Evaluates adherence to internal compliance SLA targets. | 95% compliance across all severity tiers |
| Remediation Velocity / Hazard Rate | Conditional probability of patching an open item during interval t | Tracks team patching capacity dynamically over time. | Monitored post-patch release cycles |
| Remediation Capacity Ratio | Volume of Closed vulnerabilities/Volume of newly disclosed vulnerabilities | Determines whether the vulnerability backlog is expanding or shrinking. | 1.0 (remediating faster than disclosure rate) |
Detection Maturity & Velocity
Detection maturity measures how advanced and effective your threat detection processes are, while velocity metrics track how quickly detection rules, responses, and remediations move through the pipeline.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Rule Density per Data Source | Total Documented & Tested Rules / Active Ingested Data Sources | Measures rule coverage depth across enterprise log sources. | 5 to 10 active rules per key data source |
| MITRE ATT&CK Mapping Depth | (Data Sources mapped to ATT&CK Techniques / Total Ingested Data Sources) * 100 | Quantifies structural framework alignment and threat coverage gaps. | 80% techniques covered on critical data sources |
| Detection Engineering Velocity | Count of new rules built, tuned, or deprecated per Sprint / Cycle | Measures engineering throughput, output rate, and maintenance capability. | Consistent sprint velocity (e.g., 5-8 stories/sprint) |
| Intel-to-Production Lead Time | Days/Hours elapsed from CTI report / Red Team finding to Production Rule Deployment | Measures agility in converting threat intelligence into active defense. | < 48 hours for critical CTI findings |
| False Negative (FN) Triage Rate | (Undetected Red/Purple Team Tests / Total Conducted Tests) * 100 | Identifies detection blind spots through empirical adversary simulation. | < 15% undetected simulations (driven to 0%) |
Detection Visibility & Baselines
Detection visibility metrics measure how much of your organization’s environment and activities are being monitored, while security baselines define the expected configuration and posture for systems and processes.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Identity Exposure Footprint | Count of active Users, Service Accounts, Groups, & Privileged Roles across Prod/QA/Dev | Establishes identity privilege baseline & identity attack surface size. | 100% visibility across all identity providers |
| Asset & Technology Footprint | Total monitored Workstations, Servers, & Mobile Devices broken down by OS | Justifies asset-to-analyst ratios and tracks growth across Mergers and Acquisitions or expansions. | 98% sensor coverage on corporate assets |
| Data Ingestion & Retention Depth | Daily Ingestion (GB/TB/day) & Hot/Searchable Log Retention (Days) | Models SIEM platform licensing costs and search query performance constraints. | 30-90 days hot log retention minimum |
| Network & Email Throughput Baseline | Ingress/Egress Throughput (GB/day), Email Volume, & New External Senders Count | Quantifies external attack exposure and network detection baseline volume. | Continuous baseline tracking per environment |
Detection Quality & Performance Metrics
These quantitative measures are to be used to evaluate how accurately, precisely, and reliably a system (e.g., a detection model, QA process) identifies and processes targets or defects.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Detection Reliability And Precision Efficiency (DRAPE) Index Score | DRAPE = f(True Positives, False Positives, weight w, noise penalty k) | Quantifies detection analytic quality into a composite score (signal vs. noise). | Score > 5 (Decent) or > 15 (Strong). Deprecate if < 0 |
| Normalized Endpoint Detection Ratio | (Total Endpoint Alerts / Total Monitored Endpoints) * 100 | Normalizes alert volume against fleet size to evaluate operational efficiency. | < 1.5 alerts per 100 endpoints / month |
| Normalized Network Detection Ratio | Total Network Alerts / (Ingress + Egress Network Throughput in GB) | Evaluates network rule noise independent of network traffic expansion. | Stable or decreasing trend over time |
| Normalized Identity Detection Ratio | Total Identity Alerts / Total Active User Accounts | Measures identity rule precision per account baseline. | Low baseline variance per 1,000 active users |
| Resource Forecasting Model | Projected Next-Year Alerts = Projected Asset Count * Current Alert Ratio | Helps you to predict future ticket volume to justify headcount & SIEM licensing prior to M&A. | Used annually for operational capacity planning |
Detection-as-Code (DeC) & CI/CD Metrics
These metrics combine software engineering discipline with threat detection, enabling repeatable, auditable, and measurable detection workflows.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Rule Test Coverage | (Production Detections with Automated CI Synthetic Unit Tests / Total Rules) * 100 | Ensures rules are systematically validated against regression before deployment. | > 85% automated test coverage |
| Build & Pipeline Pass Rate | (Successful CI Integration Runs / Total Pull Requests Triggered) * 100 | Measures code health, syntax accuracy, and schema compliance in detection repos. | > 90% first-pass rate |
| Detection Drift Rate | Count of out-of-band manual rule edits executed directly in SIEM/SOAR UI | Measures adherence to version-controlled Detection-as-Code deployment pipelines. | 0 manual UI edits (100% via Git) |
| Pull Request Lead Time | Median hours from Detection PR creation to Peer Review approval & Production deployment | Measures code review efficiency and operational pipeline throughput. | < 24 hours median lead time |
Threat Hunting & Capacity Metrics
These metrics help demonstrate risk reduction, improve detection, and justify resources.
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Hunt-to-Detection Conversion Rate | (Completed Threat Hunts Yielding Production Rules / Total Completed Hunts) * 100 | Evaluates efficacy of proactive threat hunting in creating durable security logic. | 20% to 30% conversion rate (Cisco PEAK framework) |
| AI Investigation Compression Ratio | Baseline Manual Hunt/Triage Duration / AI-Augmented Investigation Execution Time | Quantifies time-savings and analytical leverage gained via local/agentic LLM tooling. | 5x to 10x investigation speedup |
| Hunt MITRE ATT&CK Coverage Delta | Net-new MITRE ATT&CK techniques or sub-techniques validated via hunting per quarter | Tracks expansion of adversary technique coverage driven by hypothesis-based hunts. | +5 to 10 new sub-techniques validated per quarter |
Incident Response & Triage Efficiency Metrics
| Metric | Calculation Method | Operational Purpose | Recommended Benchmark |
|---|---|---|---|
| Mean Time to Validate (MTTV) | Mean duration from initial alert trigger to analyst confirmation (True Positive vs. Benign) | Measures initial triage efficiency before deep-dive incident response. | < 10% human override rate for low-risk actions |
| Tier-1 Escalation Accuracy | (Tier-1 Escalations Confirmed as Incidents by Tier-2/IR / Total Escalations) * 100 | Evaluates Tier-1 triage accuracy and prevents IR team fatigue. | < 15 minutes median validation time |
| Playbook Automation Coverage | (High-Severity Alert Types with Automated Playbook Actions / Total High-Sev Alert Types) * 100 | Measures degree of automated containment (e.g., host isolation, key revocation). | > 80% confirmed escalation precision |