In the realm of cybersecurity operations, speed and accuracy are inseparable. A security alert only creates value when the SOC can validate it, understand its business impact, and contain the incident before risk expands.
Yet many Security Operations Centres measure only the visible outcomes: mean time to detect (MTTD), mean time to respond (MTTR), alert volume and containment time: without examining how work actually flows between detection, triage, investigation and response.
That is where value stream mapping becomes powerful. It exposes the queues, rework, handoffs and approval delays hidden behind the dashboard. Used with DMAIC, value stream mapping helps SOC managers and IT operations leaders redesign the flow from alert raised to incident contained, while reducing unnecessary cognitive load on analysts.
The following worked example is illustrative, but the method can be applied to managed security service providers, internal SOCs, cloud operations teams and IT service environments.
Define the Scope: Map One Security Value Stream
A SOC may operate dozens of related processes. Do not begin by mapping “cybersecurity operations” as a whole. Select one measurable value stream with a clear customer and outcome.
For this guide, the scope is:
SIEM or EDR alert generated → triaged → enriched → escalated where required → incident contained and documented.
The customer is the organisation relying on the SOC to reduce exposure. The critical-to-quality requirements are:
- High-severity alerts reviewed within a defined service level.
- Accurate classification with a controlled false-positive rate.
- Complete escalation information.
- Rapid containment with minimal rework.
- Sustainable analyst workload.
This scope aligns with the practical alert triage stages described by Rapid7 and can be structured through a Lean Six Sigma project charter. A Lean Six Sigma practitioner guide can help formalise the Define and Measure phases.
Current-State Value Stream Map: Where the Queue Forms
The current-state map should be built with L1 analysts, L2 investigators, incident responders, detection engineers and the process owner. Observe the work directly rather than relying only on standard operating procedures.
A typical current-state flow is:
Alert generated → duplicate suppression → L1 triage → manual enrichment → escalation decision → L2 investigation → approval → containment → documentation and rule-tuning feedback
Mark each step with:
- Process time.
- Waiting time.
- Queue depth.
- First-pass yield.
- Handoff owner.
- Rework loop.
- System used.
- Decision criteria.

Worked Example: A 2,400-Alert Daily SOC
Consider a 24-hour SOC supporting 18,000 endpoints.
| Metric | Current baseline |
|---|---|
| Raw alerts per day | 2,400 |
| Duplicate or correlated alerts | 600 |
| Unique alert records | 1,800 |
| False-positive rate | 82% |
| L1 analysts | 12 |
| Productive capacity per analyst | 420 minutes/day |
| Average initial triage touch time | 2.8 minutes |
| Average queue depth | 780 alerts |
| MTTD | 48 minutes |
| Mean time to triage | 52 minutes |
| Daily escalations | 220 |
| L2 investigation time | 32 minutes |
| Playbook adherence | 61% |
| Mean time to respond | 6 hours 20 minutes |
| Median containment time | 94 minutes |
The capacity calculation appears close to balanced:
- Available L1 capacity: 12 × 420 = 5,040 minutes per day
- Initial triage demand: 1,800 × 2.8 = 5,040 minutes per day
However, this calculation excludes exception handling, repeated enrichment, shift handoffs, rework, coaching, documentation and escalations. The result is a process operating at theoretical capacity with no resilience.
For L2:
- Daily investigation demand: 220 × 32 = 7,040 minutes
- Available L2 capacity: six investigators × 420 = 2,520 minutes
This is the principal bottleneck. The L2 queue increases because analysts investigate too many low-value escalations, while high-risk incidents wait for specialist attention.
Analyse the Eight DOWNTIME Wastes
In the Analyse phase, combine visual tools with statistical evidence. An Affinity Diagram can organise analyst observations into meaningful categories based on natural relationships. A Pareto chart can then identify which alert rules, systems or handoffs create the greatest burden.
The eight DOWNTIME wastes appear as follows:
- Defects: Incorrect severity, incomplete escalation notes or misclassified true positives.
- Overproduction: Duplicate alerts, repeated notifications and multiple tickets for one event.
- Waiting: Alerts waiting for L1 review, L2 capacity, manager approval or infrastructure access.
- Non-utilised talent: Experienced analysts spending most of their time copying indicators between tools.
- Transportation: Moving evidence between SIEM, EDR, ticketing and threat-intelligence platforms.
- Inventory: Work in process, represented by open alerts and investigations awaiting action.
- Motion: Excessive switching between dashboards, browser tabs and communication channels.
- Extra-processing: Repeating enrichment, rewriting the same incident summary and obtaining unnecessary approvals.
The Voice of the Customer may require rapid containment. The Voice of the Business may prioritise risk reduction and predictable staffing. The Voice of the Process comes from timestamps, queue depth and defect data. These voices should be reconciled rather than optimised separately.
For example, an ANOVA can compare average triage times across day, evening and overnight shifts. Before relying on that comparison, Bartlett’s Test can assess whether group variances are sufficiently equal. A box plot can reveal skewness and outliers, while an average (mean) provides a baseline for process performance.
Use attribute data, such as Pass/Fail for playbook adherence, alongside continuous data such as triage minutes. A Z-score can highlight unusually long investigations across different alert categories. An X-bar chart can monitor average triage time alongside an R chart for variation.
The analytical logic is expressed as Y = f(x): containment performance is the outcome, while detection quality, enrichment completeness, prioritisation rules, analyst skill and approval delays are influential inputs.
Build the Future-State Flow
The future state should not simply add more analysts. It should reduce demand, shorten the path to a decision and reserve specialist capacity for confirmed risk.
The redesigned flow is:
Alert generated → automated correlation and enrichment → risk-based prioritisation → standardised L1 triage → clear escalation threshold → L2 investigation → containment authority → closed-loop learning
Key changes include:
- Autonomation (Jidoka): Automated controls detect missing context, failed enrichment or unusual alert patterns and signal the issue in real time.
- Andon signalling: A visual signal identifies critical queue growth, breached service levels or blocked containment decisions.
- Standard work: A concise playbook defines the required evidence, decision categories and escalation criteria.
- Approval redesign: Approval remains for high-risk actions, but routine containment actions use pre-authorised rules. Governance protects the organisation; excessive approval checkpoints create bottlenecks.
- Agile delivery: Short improvement iterations allow the SOC to test a detection change, measure results and refine the playbook without waiting for a large transformation programme.
- WIP limits: The team limits concurrent investigations so analysts finish priority work before starting additional cases.

Current State vs Future State
| Measure | Current state | Future state | Improvement |
|---|---|---|---|
| Raw alerts per day | 2,400 | 2,400 | Demand made visible |
| Unique records routed to people | 1,800 | 950 | 47% reduction |
| False-positive rate | 82% | 38% | 44-point reduction |
| Mean time to triage | 52 min | 14 min | 73% faster |
| MTTD | 48 min | 16 min | 67% faster |
| Daily escalations | 220 | 95 | 57% reduction |
| L2 investigation time | 32 min | 18 min | 44% faster |
| Queue depth | 780 | 120 | 85% reduction |
| Playbook adherence | 61% | 93% | 32-point increase |
| Median containment time | 94 min | 28 min | 70% faster |
| MTTR | 6 hr 20 min | 2 hr 10 min | 66% faster |
With 950 human-routed records at an average of 2 minutes each, L1 demand becomes 1,900 minutes per day, leaving capacity for exception handling and coaching. L2 demand becomes 95 × 18 = 1,710 minutes, which fits within the available 2,520 minutes.
This is the core value stream mapping insight: improvement comes from changing the relationship between demand, capacity, prioritisation and flow: not from asking analysts to work faster indefinitely.
Sequence the Kaizen Events
Kaizen should be sequenced according to dependency and risk.
Kaizen 1: Stabilise the Data and Definitions
- Agree on alert, incident, escalation and containment definitions.
- Validate timestamps across SIEM, SOAR and ticketing tools.
- Create a baseline for MTTD, MTTR, false positives and queue depth.
- Confirm the measurement system with an attribute-agreement review.
Kaizen 2: Reduce the Top Five Sources of Noise
- Build a Pareto chart of false positives by detection rule.
- Tune thresholds and suppression logic.
- Test changes on historical data before deployment.
- Track first-pass yield and unintended missed detections.
Kaizen 3: Standardise L1 Triage
- Create a timeboxed triage checklist.
- Define True Positive, Benign True Positive, False Positive and Indeterminate outcomes.
- Set escalation rules for critical assets, corroborating indicators and suspicious behaviour.
- Raise playbook adherence from 61% toward 90% or higher.
Kaizen 4: Automate Enrichment and Containment Preparation
- Attach asset criticality, user role, historical activity and threat-intelligence matches automatically.
- Introduce Jidoka checks for incomplete evidence.
- Pre-authorise low-risk containment actions.
- Use an Andon-style visual signal for blocked or overdue cases.
Kaizen 5: Control and Sustain
- Monitor queue depth, MTTD, MTTR, false-positive rate and containment time weekly.
- Use control charts for average triage time and escalation accuracy.
- Review detection rules monthly.
- Embed the new process in standard work, training and governance reviews.
A simple business case can quantify the benefit. If the redesigned flow releases 3,000 analyst minutes per day and the loaded analyst cost is $70 per hour, the theoretical capacity released is approximately $3,500 per day. Compare that value with automation and engineering costs using a business case calculator. Break-even analysis can then establish when the improvement investment is recovered.
Make Value Stream Mapping a Core SOC Capability
Value stream mapping gives cybersecurity leaders a shared view of how risk moves through the organisation. It connects technical performance with Lean concepts such as value, throughput, takt time, bottlenecks, waiting, variation and flow.
A White Belt can help the team understand the language. A Yellow Belt can support data collection and kaizen activity. A Green Belt can lead the DMAIC project, while a Black Belt can manage complex cross-functional improvement and mentor the team.
At Lean 6 Sigma Hub, our CSSC-accredited Green Belt training uses practical tools, worked examples, case studies and self-paced learning to help professionals apply improvement methods in real operating environments.
Build the capability to map your SOC value stream, prove root causes and lead measurable improvements by pursuing Lean Six Sigma certification today.
Kaizen. Kai-Care. Kai-Done. ( Lean Six Sigma)







