Value Stream Mapping for Rail Signalling and Network Control Maintenance: From Fault Alarm to Service Restored Without the Access Wait

1. Why Value Stream Mapping matters in rail maintenance

In the realm of rail signalling and network control, service disruption time is rarely created by the repair itself. The asset may require only 70 minutes of hands-on work, yet trains remain affected for many hours because the alarm waits for acknowledgement, the work order waits for planning, the crew travels, and the maintenance team waits for safe access.

Value Stream Mapping (VSM) reveals how these delays combine from fault occurrence to verified service restoration. It connects the information flow in the network control centre with the physical flow of people, tools, spares, permits and protection arrangements.

The central question is not simply, “How long does the technician take to repair the point machine?” It is:

How long does the customer-facing disruption last, and which queues create the largest share of that time?

This distinction matters because rail assets are often not the constraint. The constraint may be an access window, a planner’s queue, incomplete fault information, crew travel, unavailable spares or repeated fault calls. A VSM makes those constraints visible without compromising the safety controls required for signalling work.

Rail control-room research emphasises the importance of alarm interpretation, logbooks, asset history, diagnosis and dispatch decisions. Similarly, signalling procedures require controlled access, safe isolation, testing, certification and formal return to service. These requirements remain essential; the improvement opportunity is to make the complete flow more reliable around them.

2. Scope selection: define the rail maintenance product family

For this worked example, the product family is:

A track-circuit or point-machine failure on an electrified commuter network.

The map begins when a service-affecting fault alarm is generated or reported and ends when:

  1. The fault is corrected or safely contained.
  2. The equipment is tested.
  3. The authorised person confirms normal service.
  4. Network control records the restoration and closes the work order.

The scope includes:

  • Network control alarm generation and acknowledgement
  • Verification, triage and priority assignment
  • Work order creation and planning
  • Crew mobilisation and travel
  • Track access and protection
  • Diagnosis, spares retrieval and repair
  • Testing, certification and service restoration
  • Administrative closure and defect learning

It excludes long-term capital replacement projects and planned renewals unless they are triggered by the corrective-maintenance event.

Takt time for the maintenance workload

Assume the network has:

  • 4 response crews
  • 6.5 available maintenance hours per crew per day
  • 22 operating days per month
  • 104 service-affecting signalling defects per month

Available monthly maintenance capacity is:

4 × 6.5 × 22 = 572 hours

The maintenance takt time is therefore:

572 hours ÷ 104 defects = 5.5 hours per defect

This is the required workload rhythm, not an automatic restoration target. It tells the maintenance system whether demand is arriving faster than available response capacity.

Current-state rail signalling maintenance flow highlighting access queues

3. Current-state map: from alarm to service restored

A sample of 120 recent faults produces the following current-state profile:

Process step Active cycle time Average waiting time
Alarm generation and controller acknowledgement 3 min 12 min
Fault verification and priority triage 12 min 25 min
Work order creation 8 min 45 min
Planner scheduling 15 min 180 min
Crew mobilisation and travel 75 min 35 min
Track access and protection 15 min 240 min
Fault diagnosis 35 min 10 min
Spares retrieval 15 min 75 min
Repair execution 70 min 10 min
Testing and certification 30 min 35 min
Service restoration and closure 10 min 20 min
Total 288 min 687 min

The average restoration lead time is therefore:

288 + 687 = 975 minutes, or 16.25 hours

The value-added wrench time (diagnosis, repair and testing) is:

35 + 70 + 30 = 135 minutes

Using the Process Cycle Efficiency formula:

Flow efficiency = Value-added time ÷ Total lead time × 100

135 ÷ 975 × 100 = 13.8%

Only 13.8% of the restoration lead time is active technical work. The remainder is waiting, movement, coordination, access, logistics or administration.

The baseline operating measures are:

  • Mean time to restore service: 16.3 hours
  • First-time-fix rate: 72%
  • Crew utilisation: 42% wrench time during the sampled response shift
  • Deferred defects backlog: 148 defects
  • Work-in-process: 37 open corrective work orders
  • Work orders per crew per shift: 1.6
  • Repeated fault calls: 22% of sampled faults generated a second call within 24 hours

The three most significant delay drivers are:

  1. Access waiting: 240 minutes per fault, representing approximately 24.6% of total restoration lead time.
  2. Crew travel: 55 minutes of average travel, or approximately 5.6% of total lead time, with additional waiting caused by poor dispatch sequencing.
  3. Repeated fault calls: 22% of faults required repeat attendance. At an average 95 minutes of additional travel, access and diagnosis per repeat call, the sample generated approximately 25 hours of avoidable response effort.

The data demonstrates why the asset itself is not necessarily the bottleneck. The point machine may be repaired in 70 minutes, while the system loses 240 minutes waiting for access.

4. The eight wastes in signalling maintenance

The DOWNTIME framework helps classify waste across the value stream.

  • Defects: Incorrect alarm classification sends the wrong specialist; incomplete testing results in a repeat failure. A failed first-time-fix rate of 28% creates additional attendance and disruption.
  • Overproduction: Duplicate work orders are created for the same alarm; technicians produce detailed reports that are not used in future fault diagnosis.
  • Waiting: Crews wait for a possession or protection authority; work orders wait in a planner’s queue for the next available access window.
  • Non-utilised talent: Experienced maintainers spend time chasing permits, locating parts or re-entering information instead of diagnosing faults. Controllers may lack a structured channel to feed recurring alarm patterns into engineering.
  • Transportation: A crew travels from a distant depot because the nearest qualified team is not visible in the dispatch system; a replacement relay is transported from a central store even though the failure family is predictable.
  • Inventory: Critical point-machine components are held centrally but not at forward stores; excessive low-use inventory occupies space while high-risk spares remain unavailable.
  • Motion: Technicians walk between trackside equipment, vehicles and stores to collect tools; repeated movement occurs because standard troubleshooting kits are incomplete.
  • Excess processing: The same fault details are entered into the control log, work-order platform, spreadsheet and maintenance report. Multiple approval checkpoints delay low-risk, repeatable work.

These wastes should be prioritised by minutes and service impact, not by visual annoyance. A Pareto analysis of the sample shows that access waiting, travel and repeat attendance account for the largest improvement opportunity.

5. Future-state design: build a faster, safer flow

The future state should preserve every required protection, testing and authorisation control while removing avoidable queues.

Future-state rail maintenance using remote diagnostics and forward spares

Condition-based and remote diagnostics

Target a 30% reduction in dispatches requiring on-site diagnosis only. Use remote point-machine status, track-circuit history, alarm recurrence and asset condition indicators to provide a fault hypothesis before dispatch.

Standard troubleshooting kits and forward spares

Target 95% availability of the top five failure-mode kits at designated forward stores. Each kit should include the approved test equipment, common replacement components, connectors, documentation and safety requirements.

Daily prioritisation and level-loading board

Target 90% of service-affecting work orders prioritised within 15 minutes. A visual board should show severity, train impact, access requirement, crew capability, location and estimated restore time. This applies Heijunka principles to maintenance demand by balancing workload across crews and windows.

Pre-approved access windows

Target a reduction in average access waiting from 240 to 90 minutes. Establish recurring access windows for high-risk assets and bundle nearby corrective and preventive tasks where safe and operationally appropriate. No access window should bypass protection rules or authorisation requirements.

Standard work for the top five failure modes

Target 90% compliance with standard troubleshooting sequences and increase first-time fix from 72% to 90%. Standard work should define alarm checks, diagnostic order, required measurements, escalation triggers, approved repair methods and test evidence.

Closed-loop defect feedback to engineering

Target 100% of repeat faults reviewed within seven days. Every completed work order should capture the failure mode, cause, component, temporary action, permanent action, test result and repeat-fault status. This creates a direct Voice of the Process feedback loop for reliability engineering.

6. Current state versus future state

Measure Current state Future-state target
Mean time to restore service 16.3 hours 8.0 hours
First-time-fix rate 72% 90%
Crew travel time 55 min 30 min
Access waiting time 240 min 90 min
Deferred defects 148 60
Work orders per crew per shift 1.6 2.4
Wrench-time percentage 42% 65%
Flow efficiency 13.8% 32%

The future state does not depend on asking technicians to work faster. It improves the system around the technician: better information, better readiness, better access planning and fewer repeat visits.

7. A 90-day Kaizen sequence

Cross-functional rail maintenance team reviewing a 90-day Kaizen board

Wave 1: Days 1–30, Establish the baseline

Owner: Maintenance improvement lead, with network control and safety representatives

Actions:

  1. Define the official start and end clocks for restoration.
  2. Sample 30–50 recent point-machine and track-circuit faults.
  3. Capture every handoff, queue, travel segment, access delay and repeat call.
  4. Create the current-state VSM and Pareto of delay minutes.
  5. Confirm the top five failure modes.

Metrics and expected results:

  • 95% timestamp completeness
  • Baseline MTRS, access waiting and first-time-fix established
  • Agreement on one operational definition of “service restored”

Wave 2: Days 31–60, Pilot the flow changes

Owner: Route maintenance manager and network control manager

Actions:

  1. Launch the daily prioritisation and level-loading board.
  2. Create forward kits for the top five failure modes.
  3. Pilot remote diagnostics for one asset family.
  4. Introduce pre-approved access windows on one route.
  5. Train crews on standard troubleshooting work.

Metrics and expected results:

  • Access waiting reduced by 25%
  • Travel time reduced by 15%
  • First-time-fix improved to 82%
  • Repeat attendance reduced by 30%

Wave 3: Days 61–90, Control and scale

Owner: Asset engineering leader and Black Belt project owner

Actions:

  1. Expand the pilot to additional routes and depots.
  2. Add repeat-fault review to the weekly engineering meeting.
  3. Place restoration, access and first-time-fix metrics on the operational dashboard.
  4. Audit standard work and forward-store replenishment.
  5. Publish the future-state map and control plan.

Metrics and expected results:

  • MTRS reduced to approximately 8–10 hours
  • Access waiting at or below 90 minutes
  • First-time-fix at 90%
  • Deferred defects reduced to 60 or fewer
  • Flow efficiency increased to at least 32%

The control plan should include run charts for restoration time, an X-bar or individuals chart where appropriate, and stratification by asset type, route, crew, fault family and time of day. Means alone can conceal access outliers, so median and 90th-percentile restoration times should also be tracked.

8. Build the capability to improve complex maintenance systems

Rail signalling maintenance is a cross-functional value stream. It requires structured problem definition, reliable measurement, root-cause analysis, waste removal, statistical thinking and disciplined control.

Develop those capabilities through CSSC-accredited Lean Six Sigma training at Lean 6 Sigma Hub. The Green Belt programme is designed for professionals who lead data-driven improvement projects, while the Black Belt programme develops advanced practitioners who lead complex projects, mentor Green Belts and drive organisational change.

You can also use the Process Cycle Efficiency Calculator to quantify the gap between active work and total lead time, and the Lean Six Sigma Concepts and Glossary to reinforce the core methods.

Enrol in CSSC-accredited Green Belt or Black Belt training today and learn how to turn hidden queues into measurable, safer and more reliable flow.

Kaizen. Kai-Care. Kai-Done. Lean Six Sigma

Related Posts