1. Why Value Stream Mapping matters in rail maintenance
In the realm of rail signalling and network control, service disruption time is rarely created by the repair itself. The asset may require only 70 minutes of hands-on work, yet trains remain affected for many hours because the alarm waits for acknowledgement, the work order waits for planning, the crew travels, and the maintenance team waits for safe access.
Value Stream Mapping (VSM) reveals how these delays combine from fault occurrence to verified service restoration. It connects the information flow in the network control centre with the physical flow of people, tools, spares, permits and protection arrangements.
The central question is not simply, “How long does the technician take to repair the point machine?” It is:
How long does the customer-facing disruption last, and which queues create the largest share of that time?
This distinction matters because rail assets are often not the constraint. The constraint may be an access window, a planner’s queue, incomplete fault information, crew travel, unavailable spares or repeated fault calls. A VSM makes those constraints visible without compromising the safety controls required for signalling work.
Rail control-room research emphasises the importance of alarm interpretation, logbooks, asset history, diagnosis and dispatch decisions. Similarly, signalling procedures require controlled access, safe isolation, testing, certification and formal return to service. These requirements remain essential; the improvement opportunity is to make the complete flow more reliable around them.
2. Scope selection: define the rail maintenance product family
For this worked example, the product family is:
A track-circuit or point-machine failure on an electrified commuter network.
The map begins when a service-affecting fault alarm is generated or reported and ends when:
- The fault is corrected or safely contained.
- The equipment is tested.
- The authorised person confirms normal service.
- Network control records the restoration and closes the work order.
The scope includes:
- Network control alarm generation and acknowledgement
- Verification, triage and priority assignment
- Work order creation and planning
- Crew mobilisation and travel
- Track access and protection
- Diagnosis, spares retrieval and repair
- Testing, certification and service restoration
- Administrative closure and defect learning
It excludes long-term capital replacement projects and planned renewals unless they are triggered by the corrective-maintenance event.
Takt time for the maintenance workload
Assume the network has:
- 4 response crews
- 6.5 available maintenance hours per crew per day
- 22 operating days per month
- 104 service-affecting signalling defects per month
Available monthly maintenance capacity is:
4 × 6.5 × 22 = 572 hours
The maintenance takt time is therefore:
572 hours ÷ 104 defects = 5.5 hours per defect
This is the required workload rhythm, not an automatic restoration target. It tells the maintenance system whether demand is arriving faster than available response capacity.

3. Current-state map: from alarm to service restored
A sample of 120 recent faults produces the following current-state profile:
| Process step | Active cycle time | Average waiting time |
|---|---|---|
| Alarm generation and controller acknowledgement | 3 min | 12 min |
| Fault verification and priority triage | 12 min | 25 min |
| Work order creation | 8 min | 45 min |
| Planner scheduling | 15 min | 180 min |
| Crew mobilisation and travel | 75 min | 35 min |
| Track access and protection | 15 min | 240 min |
| Fault diagnosis | 35 min | 10 min |
| Spares retrieval | 15 min | 75 min |
| Repair execution | 70 min | 10 min |
| Testing and certification | 30 min | 35 min |
| Service restoration and closure | 10 min | 20 min |
| Total | 288 min | 687 min |
The average restoration lead time is therefore:
288 + 687 = 975 minutes, or 16.25 hours
The value-added wrench time (diagnosis, repair and testing) is:
35 + 70 + 30 = 135 minutes
Using the Process Cycle Efficiency formula:
Flow efficiency = Value-added time ÷ Total lead time × 100
135 ÷ 975 × 100 = 13.8%
Only 13.8% of the restoration lead time is active technical work. The remainder is waiting, movement, coordination, access, logistics or administration.
The baseline operating measures are:
- Mean time to restore service: 16.3 hours
- First-time-fix rate: 72%
- Crew utilisation: 42% wrench time during the sampled response shift
- Deferred defects backlog: 148 defects
- Work-in-process: 37 open corrective work orders
- Work orders per crew per shift: 1.6
- Repeated fault calls: 22% of sampled faults generated a second call within 24 hours
The three most significant delay drivers are:
- Access waiting: 240 minutes per fault, representing approximately 24.6% of total restoration lead time.
- Crew travel: 55 minutes of average travel, or approximately 5.6% of total lead time, with additional waiting caused by poor dispatch sequencing.
- Repeated fault calls: 22% of faults required repeat attendance. At an average 95 minutes of additional travel, access and diagnosis per repeat call, the sample generated approximately 25 hours of avoidable response effort.
The data demonstrates why the asset itself is not necessarily the bottleneck. The point machine may be repaired in 70 minutes, while the system loses 240 minutes waiting for access.
4. The eight wastes in signalling maintenance
The DOWNTIME framework helps classify waste across the value stream.
- Defects: Incorrect alarm classification sends the wrong specialist; incomplete testing results in a repeat failure. A failed first-time-fix rate of 28% creates additional attendance and disruption.
- Overproduction: Duplicate work orders are created for the same alarm; technicians produce detailed reports that are not used in future fault diagnosis.
- Waiting: Crews wait for a possession or protection authority; work orders wait in a planner’s queue for the next available access window.
- Non-utilised talent: Experienced maintainers spend time chasing permits, locating parts or re-entering information instead of diagnosing faults. Controllers may lack a structured channel to feed recurring alarm patterns into engineering.
- Transportation: A crew travels from a distant depot because the nearest qualified team is not visible in the dispatch system; a replacement relay is transported from a central store even though the failure family is predictable.
- Inventory: Critical point-machine components are held centrally but not at forward stores; excessive low-use inventory occupies space while high-risk spares remain unavailable.
- Motion: Technicians walk between trackside equipment, vehicles and stores to collect tools; repeated movement occurs because standard troubleshooting kits are incomplete.
- Excess processing: The same fault details are entered into the control log, work-order platform, spreadsheet and maintenance report. Multiple approval checkpoints delay low-risk, repeatable work.
These wastes should be prioritised by minutes and service impact, not by visual annoyance. A Pareto analysis of the sample shows that access waiting, travel and repeat attendance account for the largest improvement opportunity.
5. Future-state design: build a faster, safer flow
The future state should preserve every required protection, testing and authorisation control while removing avoidable queues.

Condition-based and remote diagnostics
Target a 30% reduction in dispatches requiring on-site diagnosis only. Use remote point-machine status, track-circuit history, alarm recurrence and asset condition indicators to provide a fault hypothesis before dispatch.
Standard troubleshooting kits and forward spares
Target 95% availability of the top five failure-mode kits at designated forward stores. Each kit should include the approved test equipment, common replacement components, connectors, documentation and safety requirements.
Daily prioritisation and level-loading board
Target 90% of service-affecting work orders prioritised within 15 minutes. A visual board should show severity, train impact, access requirement, crew capability, location and estimated restore time. This applies Heijunka principles to maintenance demand by balancing workload across crews and windows.
Pre-approved access windows
Target a reduction in average access waiting from 240 to 90 minutes. Establish recurring access windows for high-risk assets and bundle nearby corrective and preventive tasks where safe and operationally appropriate. No access window should bypass protection rules or authorisation requirements.
Standard work for the top five failure modes
Target 90% compliance with standard troubleshooting sequences and increase first-time fix from 72% to 90%. Standard work should define alarm checks, diagnostic order, required measurements, escalation triggers, approved repair methods and test evidence.
Closed-loop defect feedback to engineering
Target 100% of repeat faults reviewed within seven days. Every completed work order should capture the failure mode, cause, component, temporary action, permanent action, test result and repeat-fault status. This creates a direct Voice of the Process feedback loop for reliability engineering.
6. Current state versus future state
| Measure | Current state | Future-state target |
|---|---|---|
| Mean time to restore service | 16.3 hours | 8.0 hours |
| First-time-fix rate | 72% | 90% |
| Crew travel time | 55 min | 30 min |
| Access waiting time | 240 min | 90 min |
| Deferred defects | 148 | 60 |
| Work orders per crew per shift | 1.6 | 2.4 |
| Wrench-time percentage | 42% | 65% |
| Flow efficiency | 13.8% | 32% |
The future state does not depend on asking technicians to work faster. It improves the system around the technician: better information, better readiness, better access planning and fewer repeat visits.
7. A 90-day Kaizen sequence

Wave 1: Days 1–30, Establish the baseline
Owner: Maintenance improvement lead, with network control and safety representatives
Actions:
- Define the official start and end clocks for restoration.
- Sample 30–50 recent point-machine and track-circuit faults.
- Capture every handoff, queue, travel segment, access delay and repeat call.
- Create the current-state VSM and Pareto of delay minutes.
- Confirm the top five failure modes.
Metrics and expected results:
- 95% timestamp completeness
- Baseline MTRS, access waiting and first-time-fix established
- Agreement on one operational definition of “service restored”
Wave 2: Days 31–60, Pilot the flow changes
Owner: Route maintenance manager and network control manager
Actions:
- Launch the daily prioritisation and level-loading board.
- Create forward kits for the top five failure modes.
- Pilot remote diagnostics for one asset family.
- Introduce pre-approved access windows on one route.
- Train crews on standard troubleshooting work.
Metrics and expected results:
- Access waiting reduced by 25%
- Travel time reduced by 15%
- First-time-fix improved to 82%
- Repeat attendance reduced by 30%
Wave 3: Days 61–90, Control and scale
Owner: Asset engineering leader and Black Belt project owner
Actions:
- Expand the pilot to additional routes and depots.
- Add repeat-fault review to the weekly engineering meeting.
- Place restoration, access and first-time-fix metrics on the operational dashboard.
- Audit standard work and forward-store replenishment.
- Publish the future-state map and control plan.
Metrics and expected results:
- MTRS reduced to approximately 8–10 hours
- Access waiting at or below 90 minutes
- First-time-fix at 90%
- Deferred defects reduced to 60 or fewer
- Flow efficiency increased to at least 32%
The control plan should include run charts for restoration time, an X-bar or individuals chart where appropriate, and stratification by asset type, route, crew, fault family and time of day. Means alone can conceal access outliers, so median and 90th-percentile restoration times should also be tracked.
8. Build the capability to improve complex maintenance systems
Rail signalling maintenance is a cross-functional value stream. It requires structured problem definition, reliable measurement, root-cause analysis, waste removal, statistical thinking and disciplined control.
Develop those capabilities through CSSC-accredited Lean Six Sigma training at Lean 6 Sigma Hub. The Green Belt programme is designed for professionals who lead data-driven improvement projects, while the Black Belt programme develops advanced practitioners who lead complex projects, mentor Green Belts and drive organisational change.
You can also use the Process Cycle Efficiency Calculator to quantify the gap between active work and total lead time, and the Lean Six Sigma Concepts and Glossary to reinforce the core methods.
Enrol in CSSC-accredited Green Belt or Black Belt training today and learn how to turn hidden queues into measurable, safer and more reliable flow.
Kaizen. Kai-Care. Kai-Done. Lean Six Sigma








