Interpretable Prediction of Geopolymer Concrete Compressive Strength Using DBO–CatBoost and SHAP Analysis
Nima Saeedi (University of Tabriz)
Zahra Mohammadipour Novin (University of Tabriz)
Amirreza Shirini (Sahand University of Technology)
Sina Samadi Gharehveran (University of Tabriz)
Siamak Pedrammehr (Tabriz Islamic Art University)
Mohammad Fotouhi (TU Delft - Civil Engineering & Geosciences)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
The construction sector faces a critical need to minimize its carbon footprint, which is currently stimulating the development of geopolymer concrete using recycled coarse aggregates as an eco-friendly material compared with Portland cement. Accurate prediction of the compressive strength of this eco-efficient concrete is complex, however, as a result of the complex, non-linear interactions between many of the mix-design and curing parameters. Although modern scientific literature and engineering practices have increasingly adopted machine learning (ML) for concrete strength prediction, a significant scientific gap remains. Most existing studies rely on “black-box” models that lack sufficient interpretability and frequently overlook the severe risk of data leakage during validation, limiting their practical engineering application. To address this gap, this study proposes a robust, data-leakage-aware framework driven by a rigorous nested GroupKFold cross-validation strategy. By grouping concrete samples by their unique Mix_ID, this approach ensures genuine generalization to entirely unseen mixtures. Within this reliable validation scheme, the CatBoost algorithm is utilized for compressive-strength prediction, with the Dung Beetle Optimizer (DBO) serving as an effective tool for hyperparameter tuning. The evaluation results across multiple random seeds show that the DBO–CatBoost model significantly outperforms the default CatBoost, rigorously tuned baseline models (Support Vector Regression and Random Forest), and a comparative metaheuristic benchmark (PSO–CatBoost). It achieves the most stable distribution of errors and excellent predictive accuracy (Test (Formula presented.), RMSE = (Formula presented.) ). In addition, the model predictions were demystified using the methods of SHapley Additive exPlanations (SHAP) and partial dependence plots (PDPs). The interpretability analysis revealed strong statistical associations, showing that Curing Time and Coarse Aggregate are the most prominent predictive features and the strongest pairwise interaction between each other; the NaOH molar concentration is the most important second-level influence on optimization of strength. Overall, the framework provides a robust data-driven screening tool that can assist in preliminary mix-design evaluation. By reducing the reliance on extensive empirical “trial and error” approaches, this predictive model supports more efficient material usage and facilitates preliminary optimization of low-carbon concrete formulations. Theoretically, this study advances the fundamental science of geopolymer materials by explicitly quantifying the complex, non-linear interactions between alkaline activators, curing conditions, and recycled aggregates. This provides a robust data-driven theoretical foundation for designing and optimizing next-generation eco-friendly concrete products and structures.