Goal of the analysis:
Quantify the absolute hours equipment is unexpectedly unavailable for production, pinpoint the dominant causes, and translate those hours into lost throughput and cost. Unplanned Downtime Hours focuses on the “breakdown/repair” and other unplanned stop states (faults, safety stops, unplanned changeover overruns, waiting for maintenance) and is the primary driver of the Availability component of OEE. Executives use this analysis to: (1) raise throughput without capex by recovering capacity at the constraint; (2) stabilize lead times and OTIF; (3) size/target TPM, condition monitoring, and spares; and (4) align scheduling, changeover discipline, and material readiness so that true mechanical downtime—not starvation or blockage—is addressed.
Data required:
- Machine state/event telemetry (MES/SCADA/OEE):
- Time-stamped states: Running, Setup/Changeover, Breakdown/Repair, Fault/Safety stop, PM, Starved, Blocked, Micro-stop, Idle; event start/stop times and reason codes.
- Policy thresholds for micro-stop vs downtime (e.g., micro-stop < 60 sec; downtime ≥ 60 sec).
- Maintenance records (CMMS):
- Corrective and preventive work orders with failure codes, timestamps, parts used, MTBF/MTTR history, backlog, technician notes.
- Planning and execution context:
- Schedules/dispatch lists, planned changeovers and durations, planned PM windows, frozen windows, labor rosters and skills.
- Flow and material readiness:
- Kit completeness, shortage logs, upstream WIP buffer levels, downstream buffer/pack status to verify starved/blocked vs true downtime.
- Sensors and environment:
- Condition data (vibration, temperature, current), alarms, utility events (air, steam, power), ambient temperature/humidity for sensitive processes.
- Throughput and financial overlays:
- Standard time per unit or demonstrated rate by SKU, FPY at the bottleneck, contribution margin per unit/hour, overtime/premium freight costs.
- Normalization/governance:
- Shift calendars, planned vs unplanned stop taxonomy, time zones, policy on classifying setup overruns, micro-stop thresholds.
Detailed step-by-step instruction on how to conduct the analysis:
- Define scope and metrics.
- Scheduled Time = shift time − breaks.
- Planned Stops = planned changeovers + planned PM (per schedule).
- Unplanned Downtime Hours = Σ durations in Breakdown/Repair, Fault/Safety stop, unplanned extensions to setup/PM, waiting maintenance; exclude Starved/Blocked and micro-stops per policy.
- Supporting: Availability = 1 − (Unplanned Downtime ÷ Planned Production Time); MTBF and MTTR.
- Extract and align datasets.
- Pull 8–12 weeks of state logs from MES/SCADA and corrective/PM work orders from CMMS; align to common machine IDs and timestamps.
- Join schedule for planned stops and WIP/buffer/kit data to validate starved/blocked classifications.
- Clean and classify events.
- Normalize reason codes; collapse duplicates/rapid oscillations; enforce micro-stop thresholding.
- Reclassify mis-tagged starvation/blockage out of downtime using WIP sensors and shortage logs.
- Flag setup/PM overrun minutes beyond plan as unplanned downtime (separate reason).
- Compute Unplanned Downtime Hours by asset and period.
- Aggregate by machine/line/shift/day/week: total Unplanned Downtime Hours, and share of Scheduled Time.
- Break down by reason family: Mechanical, Electrical/Controls, Utilities, Safety/Quality hold, Setup overrun, Waiting maintenance, Other.
- Reliability analysis (MTBF/MTTR).
- Derive failure events (transitions to Breakdown/Repair) to compute MTBF and MTTR per asset; plot control charts and distributions.
- Identify chronic offenders: repeat failure modes/components with high downtime minutes per event or high frequency.
- Pareto and loss bridges.
- Create Pareto of Unplanned Downtime Hours by reason, by asset, and by shift/crew.
- Minutes bridge: Planned Production Time → −Planned stops → −Unplanned downtime (by cause) → Operating Time; overlay micro-stop time for context.
- Translate hours to throughput and cost.
- Lost units = Unplanned Downtime minutes ÷ Standard minutes per good unit (or × demonstrated rate × FPY).
- Capacity days lost at constraint = Unplanned Downtime hours at constraint ÷ Effective daily capacity hours.
- Monetize: Lost contribution margin + overtime/premium freight triggered + incremental maintenance/parts.
- Segment and localize drivers.
- Slice by SKU family and changeover pair (to expose overrun patterns), before/after PM windows, weekdays vs weekends, environment/utility excursions.
- Compare shifts/crews to find capability gaps; correlate with kit completeness and buffer health to rule out flow-induced stops.
- Scenario and what-if.
- Model impact of TPM (MTBF +X%), spares kitting (MTTR −Y%), condition monitoring on top subsystems, SMED (setup overrun −30–50%), PM rescheduling to demand valleys, and kitting/buffer sizing (starved/blocked reduction) on Unplanned Downtime Hours and throughput.
- Estimate ROI and expected clearance of backlog days.
- Integrity checks.
- Ensure states sum to Scheduled Time; micro-stops are not counted as downtime; planned vs unplanned stop rules consistently applied.
- Spot-check events against CMMS work orders and operator logs; audit reason-code hygiene (limit “unknown”).
Format of the output of analysis:
- Executive scorecard: Unplanned Downtime Hours (weekly/shift) by asset and plant, Availability %, MTBF/MTTR, top 5 causes, lost units and capacity days, $ impact, trend vs target.
- Pareto charts: downtime hours by cause, by asset, by shift/crew; repeat offender components.
- Time-state visuals: stacked bars per shift/day (Run, Setup, PM, Breakdown, Fault, Starved, Blocked, Micro-stops); event Gantt timelines.
- Reliability panel: MTBF/MTTR control charts; early-life vs wear-out patterns for top assets.
- Heatmaps: Unplanned Downtime Hours by asset × shift and by SKU family/changeover pair.
- Scenario/ROI deck: projected hour reductions, availability lift, units recovered, and $ benefit from TPM, PdM, SMED, PM moves, spares, and flow fixes.
How to interpret results:
- High total hours with low MTBF: Frequent failures—focus on reliability fundamentals (TPM, RCAs on repeat modes, condition monitoring) and component quality.
- High hours with high MTTR: Few but long repairs—improve serviceability: spares kitting, troubleshooting SOPs, technician coverage, and access/ergonomics.
- Large “setup overrun” share: Changeover design/discipline gap—run SMED, externalize and parallelize tasks, tighten planned times, and sequence by family.
- Spikes aligned to specific SKUs/pairs: Recipe/sizing transitions causing faults—centerline parameters, verify recipes, and adjust sequence rules.
- Downtime hours flagged but buffer/starved evidence present: Misclassification—address material readiness (kitting/ATP) and release control; don’t misdirect reliability resources.
- Shift-to-shift disparities: Capability or supervision issue—standardize operator care, cross-train, and coach troubleshooting.
Steps a company can take to improve on this measure:
- Reliability engineering and TPM:
- Implement autonomous maintenance (clean/inspect/lubricate) with daily checklists; defect tagging with rapid response.
- RCA the top 3 failure modes; update PFMEA/control plans; install condition monitoring (vibration/temp/current) on critical subsystems.
- Optimize PM using RCM; move PMs to demand valleys and time-box tasks to reduce overruns.
- Maintenance execution and spares:
- Pre-kit critical spares; maintain min/max and lead times; standardize breakdown response and diagnostics; improve wrench time via CMMS planning/kitting.
- Train technicians on top failure diagnostics; maintain escalation protocols.
- Changeover stability (to curb overruns/faults):
- Run SMED on top “from→to” pairs; quick-release tooling, presetting; centerline sheets with digital recipe verification; first-article checks to avoid fault-induced downtime.
- Flow and material readiness:
- Kitting/line-side supermarkets with ≥98% completeness; shortage dashboards; drum-buffer-rope or CONWIP to prevent starvation/blockage misclassification.
- People and governance:
- Daily tiered reviews of downtime hours and Pareto; publish MTBF/MTTR by asset; assign owners and due dates; audit reason-code accuracy.
- Certify operators in operator care and basic troubleshooting; align incentives to availability at the constraint and OTIF.
- Digital enablement:
- Automate event capture; add edge analytics for early anomaly detection; integrate MES–CMMS for closed-loop work orders triggered by alarms.
- Example scenarios:
- If Unplanned Downtime is 32 hours/week on the filler (MTBF 16h, MTTR 80m), install vibration monitoring on the drive train, kit spares, and move PMs to weekends; target ≤18 hours/week and MTTR ≤45m in 8 weeks.
- If 25% of hours are setup overruns on three pairs, execute SMED and enforce 48h frozen window; expect −40–50% overrun hours and +8–10% throughput at the constraint.
- If one shift contributes 2× downtime vs others, deploy centerline audits, troubleshooting training, and senior tech coverage; cut shift gap by half within two cycles.
Benchmark comparisons:
General benchmarks (directional):
- Discrete/high-mix lines: unplanned downtime typically 5–10% of Planned Production Time; best-in-class 2–5%. For a 24h day, this equates to ~1–2.4 hours/day/asset at world-class levels.
- Process/continuous: unplanned downtime often ≤2–5% (≈0.5–1.2 hours/day/asset on 24h runs); events are fewer but longer—focus on MTTR.
- Setup overrun share after SMED: ≤10–15% of total stop time (planned + unplanned).
- MTBF/MTTR directional: discrete MTBF 30–100h with MTTR 20–60m; process MTBF days with MTTR <60m on critical assets.
Constructing internal benchmarks:
- Track weekly Unplanned Downtime Hours by asset/shift with quartiles; set alerts (e.g., >8 hours/asset/week or 2 consecutive weeks trending up).
- Publish MTBF/MTTR and top-cause Pareto per asset; require RCA closure on repeat modes; roll out condition monitoring for assets above thresholds.
- Pair downtime hours with lost units, backlog days, and OTIF to prioritize the constraint and high-value assets.
- Rebaseline after SMED/TPM, major maintenance, or mix shifts; codify best-performing cells’ centerlines, PM timing, and spares kits as standards.