Expert-Enriched Predictive Modeling for Engine Test Cell Outcomes in Low-Data MRO Settings

A Scenario-Based Study of Expert Knowledge Integration and Data Quality Impact

Master Thesis (2025)
Author(s)

K. van Beek (TU Delft - Technology, Policy and Management)

Contributor(s)

J.A. Annema – Mentor (TU Delft - Technology, Policy and Management)

S. Balakrishnan – Graduation committee member (TU Delft - Technology, Policy and Management)

A.C. Smit – Graduation committee member (TU Delft - Technology, Policy and Management)

Faculty
Technology, Policy and Management
More Info
expand_more
Publication Year
2025
Language
English
Coordinates
52.3105, 4.7683
Graduation Date
19-08-2025
Awarding Institution
Delft University of Technology
Programme
Management of Technology (MoT)
Faculty
Technology, Policy and Management
Downloads counter
135
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Engine test runs are among the most resource-intensive steps in the aircraft maintenance cycle. On average, each run consumes around 13,000 liters of fuel, sometimes peaking at nearly 64,000 liters, equivalent to more than 200 tonnes of CO2 emissions per test. Beyond the environmental cost, a failed test can delay engine delivery by up to three months, as the engine often has to re-enter the full maintenance cycle. These failures not only drive up costs and resource consumption, but also disrupt planning and reduce the availability of engines for airline operations. In this context, the ability to predict rejection risk before a test is performed could unlock substantial savings, improve reliability, and reduce the environmental footprint of engine maintenance.

This thesis explores the development and possible implementation of a predictive model that assesses the risk of engine rejection during test cell acceptance runs at an aviation Maintenance, Repair, and Overhaul (MRO) facility. Engine test runs are costly, resource-intensive, and critical for quality assurance, but occasional failures lead to significant delays, rework, and planning disruptions. Early identification of engines at high risk of rejection could enable proactive planning and improve operational efficiency.

The study focuses on the challenges of predictive modeling under real-world constraints: a small, imbalanced dataset and limited failure labels. To overcome these limitations, the research integrates domain expert knowledge at multiple stages of the modeling pipeline, including feature selection, data curation, and the specification of Bayesian priors. A hybrid methodology combining CRISP-DM and the Design Science Research Methodology (DSRM) guides the iterative model development, evaluation, and implementation process.

Various modeling techniques were explored, including logistic regression, random forests, and Bayesian logistic regression with informative priors. The study compares conventional ML models with an expert-informed alternative under different performance objectives. These different performance objectives are prioritizing recall to support early intervention, and later precision to support robust planning. Model evaluation was performed through 5-fold cross-validation.

Results show that expert knowledge significantly improves both the quality of input data and the relevance of engineered features. While Bayesian priors didn’t really contribute to performance, the most impactful improvements stemmed from expert-driven data cleaning and categorical encoding. However, model interpretability remained limited: although rejections could be predicted, the model could not reliably indicate why they occurred.

Consequently, the model’s use case was reframed from an initial ’proactive technical intervention’ to ’planning support’. Under a precision-tuned configuration, the model achieved 50% precision in flagging high-risk engines. Because 87% of rejections were associated with long turnaround delays, each high-precision flag became a useful proxy for high-impact disruptions. This insight provides valuable input for planners seeking to reduce schedule volatility.

The study contributes to literature on expert-informed learning by highlighting how domain expertise can improve data quality, and not just simple model design, in low-data environments. It also offers an empirical evaluation of how different types of expert input affect predictive performance. Practically, the study provides a roadmap for applying expert-informed modeling in safety-critical, data-constrained MRO settings.

Recommendations for future research include validating the approach on other engine types, improving expert elicitation methods by reducing the subjective nature, exploring scalable data curation (e.g., via automation or LLMs), and developing interpretable models that can suggest targeted maintenance actions.

Files

License info not available
warning

File under embargo until 19-08-2026