Inadequate railway drainage can trigger flooding, accelerate track degradation, and ultimately cause the sudden failure of tracks, slopes, and embankments. Routine visual inspections score each drainage asset’s service-level on a scale of one to five, yet existing prediction studies typically assume a linear mapping between readily measured drainage-asset characteristics (e.g., pipe material, shape, slope, and local rainfall) and those ordinal service-level grades. Such linearity greatly oversimplifies the nonlinear, multifactor deterioration processes that govern drainage performance. This study introduce a risk-based redefinition of service level based on the direction and magnitude of change between consecutive inspections (low, medium, or high risk). Using this reformulated target, we build and compare three supervised machine learning models, namely, multinomial logistic regression, random forests, and artificial neural networks, to predict the drainage asset’s future risk category from its physical, environmental, and operational attributes. The modeling framework explicitly tackles practical data challenges: misclassification bias in visual grades, repeated measurements of the same asset, severe class imbalance, and the categorical nature of most variables. The approach is demonstrated on two UK main-line routes (London–Cardiff and Edinburgh–Glasgow) comprising approximately 10,000 drainage assets. After careful oversampling, random forests delivered the highest recall for high-risk assets, outperforming the linear baseline and the other nonlinear model. Results confirm that abandoning the linearity assumption, redefining risk to capture service-level dynamics and applying tailored preprocessing markedly improved predictive accuracy and interpretability. This study therefore provides railway owners with a data-driven tool to prioritize drainage maintenance and reduce network disruption.
Publications
Machine Learning Approach to Redefining Risk in Railway Drainage Systems