Circular Image

D.P. Solomatine

info

Please Note

86 records found

Journal article (2026) - Vitali Diaz, Ahmed A. Osman, Gerald A. Corzo Perez, Shreedhar Maskey, D.P. Solomatine
More severe and prolonged droughts observed in recent decades require improved methods to predict impacts on agriculture. Crop-growth models estimate yield and plant development variables and are widely used to assess drought impacts; however, they are not explicit forecasting tools, as their accuracy is constrained by physical assumptions, data availability, and multiple sources of uncertainty. To address these limitations, machine learning (ML) models have been increasingly applied for crop yield prediction, typically using drought indices as input, while spatial drought characteristics remain underexplored. This research develops an ML framework that incorporates the spatial extent of drought to predict seasonal crop yield. The framework combines artificial neural network (ANN) and polynomial regression (PR) models, with PR providing baseline estimates and ANN delivering refined predictions. The approach was tested using 50 years of historical crop yield data and drought areas derived from the Standardised Precipitation Evapotranspiration Index at multiple aggregation periods (1–12 months). Results show ANN models consistently outperform PR models, achieving lower prediction errors, with root mean square error values as low as 48.1 kg/ha in best-performing cases. The results demonstrate that spatiotemporal drought area dynamics and their temporal aggregation provide an effective preprocessing strategy for ML-based drought impact prediction. ...
Journal article (2026) - Ja Ho Koo, Edo Abraham, Andreja Jonoski, Dimitri P. Solomatine
Dealing with uncertainty in predicted inflows presents a major challenge in optimal reservoir flood control. Scenario-based stochastic control approaches address this by generating multiple inflow time series from probabilistic models, each representing a possible future with associated likelihoods. However, using too many scenarios increases computational complexity, while too few may compromise representativeness. Although the two critical steps of scenario generation and reduction have been extensively explored in other fields, their application to reservoir inflow dynamics remains limited. This study develops and applies a probabilistic data-driven model, specifically, a Bayesian Neural Network (BNN), for scenario generation. While the model exhibits limitations in predicting peak inflows due to data scarcity, it effectively captures temporal dependencies in inflow time series and achieves high short-term accuracy, as measured by the Nash–Sutcliffe Efficiency Coefficient (NSE) and Root Mean Squared Error (RMSE), though performance declines over longer horizons. For scenario reduction, four distance measures widely used in other domains, i.e., the Manhattan, Euclidean, Wasserstein, and energy distances, are evaluated. Experimental results show that the energy distance best preserves the statistical properties of the full scenario set, followed by the Manhattan and Euclidean distances. However, in terms of retaining extreme inflow scenarios, which are critical for flood control, the Manhattan and Euclidean distances outperform others based on a custom index measuring the envelope size of the original scenario set using the l1-norm. In terms of computational efficiency of scenario reduction approaches, the energy distance is the most expensive (quadratic in m, the number of reduced scenarios), while the Wasserstein scales linearly. In the examples used, reduced sets are shown to adequately capture extremes when the number of scenarios m≥30. Considering the trade-off between preserving extremes and computational cost, the Manhattan and Euclidean distances with m=30 are recommended as a practical choice for reservoir inflow scenario reduction. ...
Journal article (2026) - Santiago Duarte, Gerald Corzo, Dimitri Solomatine, Remko Uijlenhoet
AbstractStudy RegionThe study region is the Magdalena River basin in Colombia. The basin was divided into three distinct regions (Andean, Caribbean, and Pacific) and analyzed across different elevations.Study FocusThe study proposes a Spatiotemporal Non-Linear Dynamics Assessment (SNLDA) framework to compare ERA5-Land reanalysis data with in-situ rain gauge observations. It specifically examines the constraints imposed by nonlinear dynamical processes and their associated space-time complexities on the representation of precipitation, particularly in a tropical region. The SNLDA framework incorporates three main components: (i) standard performance metrics (e.g., correlations, RMSE, and dry spell duration), (ii) rainfall spatiotemporal objects (characterizing precipitation events through attributes such as volumes and start-end centroids), and (iii) non-linear dynamics complexity (reconstructing dynamical behavior from time series and evaluating attractors properties, including the Hurst and Lyapunov exponents). These elements were analyzed both individually and in combination. Daily ERA5-Land information (0.1°x0.1°) and in-situ rain gauge data comprising 558 stations from 1980 to 2020 were used, enriched by an Inverse Distance Weighting (IDW) interpolation (0.1°x0.1°) to facilitate comparison across spatial scales.New Hydrological Insights for the RegionOverall, ERA5-Land overestimates precipitation, producing shorter, more frequent events while poorly representing extreme wet and dry spells.Andean region: ERA5-Land overestimates rainfall, with largest errors at low elevations, driven by unresolved spatiotemporal object volumes displacements and nonlinear processes.Caribbean region: ERA5-Land shows the highest errors in nonlinear dynamics and extremes, despite lower annual bias and RMSE.Pacific region: ERA5-Land strongly overestimates precipitation volumes and RMSE, while nonlinear errors remain low; these biases are mainly driven by spatiotemporal objects displacement. ...

Spatial equilibrium-based water resources allocation in the upper and middle reaches of the Huaihe River Basin

Journal article (2026) - Jitao Zhang, Jinyu Meng, Zengchuan Dong, Dimitri Solomatine, Hui Xu, Wenzhuo Wang, Daoli Wang, Tianyan Zhang, Guang Yang
Optimized water resource allocation is critical for promoting sustainable regional development. However, the intensification of climate change impacts has increased the variability of water availability, thereby compounding uncertainty and complexity in water allocation decision-making. Meanwhile, mismatches in the spatial distribution of water supply and demand have also become an essential constraint to effectiveness and fairness of water allocation. Prior studies show weak correspondence between modeling frameworks to real world conditions and insufficiently describe regional balance, as well as the dynamic interaction between water inflow and demand under varying hydrological regimes. To address these challenges, this study develops a multi-objective water resource allocation model guided by spatial equilibrium principles, ensuring not only fair inter-regional allocation but also balanced co-development among water, socio-economic, and ecological subsystems. Regional balance is quantified with the Gini coefficient. Subsystem co-development is enforced via the coupling coordination degree. To address uncertainty, a nested multi-scenario robust optimization framework is proposed. It applies copula function to generate realistic encounter scenarios and applies a multi-objective probabilistic robustness evaluation to test solution stability and adaptability across scenarios. Applied to the upper and middle reaches of the Huaihe River Basin (HRB), the approach reduces the water-deficit rate by 45.42 % under extreme dry scenarios, lifts the coordination index to > 0.75, and increases reliability by 45.3 % compared with a traditional robust optimization baseline, demonstrating effectively optimize complex water resource systems while keeping both fairness and stability. This research offers a theoretical and practical framework for optimizing water resource allocation in complex and uncertain environments, contributing to the advancement of resilient and equitable water governance. ...
Journal article (2025) - Ja-Ho Koo, Edo Abraham, Andreja Jonoski, Dimitri P. Solomatine
Model predictive control (MPC) is an optimal control strategy suited for flood control of water resources infrastructure. Despite many studies on reservoir flood control and their theoretical contribution, optimisation methodologies have not been widely applied in real-time operation due to disparities between research assumptions and practical requirements. To address this gap, we include practical objectives, such as minimising the magnitude and frequency of changes in the existing outflow schedule. Incorporating these objectives transforms the problem into a multi-objective nonlinear optimisation problem that is difficult to solve in real time. Additionally, it is reasonable to assume that the weights and some parameters, considered the operators’ preferences, vary depending on the system state. To overcome these limitations, we propose a framework that converts the original intractable problem into parameterised linear MPC problems with dynamic optimisation of weights and parameters. This is done by introducing a model-based learning concept. We refer to this framework as Parameterised Dynamic MPC (PD-MPC). The effectiveness of this framework is demonstrated through a numerical experiment for the Daecheong multipurpose reservoir in South Korea. We find that PD-MPC outperforms standard MPC-based designs without a dynamic optimisation process for the objective weights and model parameter. Moreover, we demonstrate that the weights and parameters vary with changing hydrological conditions. ...

Learning-based explicit and switched model predictive control approaches

Journal article (2025) - Ja-Ho Koo, Ali Moradvandi, Edo Abraham, Andreja Jonoski, Dmitri P. Solomatine
Effective reservoir flood control demands real-time decision-making that balances multiple objectives. However, traditional optimization approaches are often too computationally intensive and become intractable when considering dynamically changing preferences of operators, modelled as weights of different objectives. This study aims to develop tractable real-time flood control strategies that maintain performance while reducing computational complexity. We propose two data-driven approaches based on Model Predictive Control (MPC): (1) an explicit MPC using deep neural networks to directly determine optimal outflow schedules, and (2) a switched MPC that produces optimal weights of objectives based on hydrological conditions. Both methods leverage offline learning from an online Parameterized Dynamic MPC framework incorporating state-dependent weights. We tested these approaches on South Korea’s Daecheong multipurpose reservoir using historical flood events with various patterns. The explicit MPC demonstrated reliable performance under conditions similar to its training data. However, it showed frequent changes in outflow schedules and constraint violations for scenarios outside training data. In contrast, the switched MPC maintained robustness across all test scenarios due to a linear optimization process in a receding horizon manner, though with slightly reduced performance compared to the explicit MPC under scenarios inside the range of training data. Most significantly, both approaches reduced computation time from approximately 10 minutes to less than one second, making real-time implementation feasible. This dramatic improvement enables prompt decision-making during rapidly evolving flood events while maintaining near-optimal control performance. ...
Journal article (2025) - Eliana Torres, Shreedhar Maskey, Gerald Corzo, Remko Uijlenhoet, Dimitri Solomatine
Study region: Madeira River basin, southwestern Amazonia Study focus: This study investigates spatial and temporal changes in precipitation, evaporation, and streamflow, and their relationship with deforestation in the Madeira River basin, the largest Amazonian sub-basin. We applied Mann-Kendall trend analysis, change-point detection, and correlation analysis across multiple spatial scales, using satellite, reanalysis, and observed data from 1981 to 2015. These methods enabled us to detect long-term trends, identify shifts, and quantify the relationships between forest loss and hydrological changes New hydrological insights for the region: The basin experienced an average deforestation rate of 2810 km² per year from 2001 to 2020, predominantly in the Brazilian portion. Between 1981 and 2016, we observed statistically significant negative trends in precipitation, evaporation, and streamflow, especially in the most deforested areas during the wet season. Correlation analysis (2001–2015) showed a statistically significant and positive relationship between forest area and evaporation in wet months (r = 0.73, p < 0.1) and a negative correlation between forest area and streamflow during the same season (r = –0.6, p < 0.1). These findings highlight the critical role of forests in modulating hydrological processes, supporting the hypothesis that deforestation may reduce evaporation, alter moisture recycling, and slow the water cycle. While our results are robust, we acknowledge that factors such as climate variability and land management practices may also influence hydrological changes and should be considered in future research. ...
Book chapter (2024) - Gerald A. Corzo Perez, Dimitri P. Solomatine
In recent years, there has been a surge of interest in machine learning (ML) and artificial intelligence (AI) due to the effectiveness of deep learning algorithms and the increasing availability of large data sets. This chapter provides a brief overview of the applications of AI and ML techniques in hydroinformatics, a field that deals with advanced information technology, data analytics, and modeling for aquatic environment management. Data-driven models are becoming more common in water management as they can reveal hidden patterns in data and offer improved accuracy in certain situations. This chapter highlights the importance of spatiotemporal data analysis, pattern recognition, and optimization approaches in water resources management under uncertainty. It does not offer a comprehensive review of all methods but rather focuses on selected ML techniques widely used in water-related problems. Additionally, the chapter discusses the challenges associated with using ML models, such as black-box criticisms, and the potential of hybrid models that combine the strengths of ML and physically based process models for more robust solutions in hydroinformatics. ...
Journal article (2024) - Boris I. Gartsman, Dimitri P. Solomatine, Tatiana S. Gubareva
Contemporary distributed hydrological models are detailed and mathematically rigorous, but their calibration and testing can be still an issue. Often it is based on the quadratic measure of the calculated and observed hydrographs proximity at one outlet gauge station, typically on the Nash-Sutcliffe model efficiency coefficient (NSE). This approach seems insufficient to calibrate a model with hundreds of spatial elements. This paper presents using a multi-dimensional estimator of modeling quality, being a natural generalization of the traditional NSE but which would aggregate data from several hydrological stations using Principal Component Analysis (PCA). The method was tested on the ECOMAG model developed for a sub-basin (24,400 km2, with 15 gauges) of the Ussuri River in Russia. The results show that the presented version of the multi-dimensional NSE with PCA in calibration of spatially-distributed hydrological models has a number of advantages compared to other methods: the reduced dimensionality without loss of important information, straightforward data analysis and the automated calibration procedure; objective separation of the deterministic signal from the noise, calibration using the “informational kernel” of data, leading to more accurate parameters’ estimates. Additionally, the introduced notion of the “compact” dataset allow to interpret physical-geographical homogeneity of the basins in mathematic manner, which can be valuable for hydrological zoning of the basins, hydrological fields analysis, and structuring the models of large basins. There is no doubt that further development and testing of the proposed methodology is advisable in solving spatial hydrological problems based on distributed models, such as managing a cascade of reservoirs, creating hydrological reanalyses, etc. ...

Case study of water resources allocation in the Huaihe River basin

Journal article (2024) - Jitao Zhang, Dmitri Solomatine, Zengchuan Dong
Water resources managers need to make decisions in a constantly changing environment because the data relating to water resources are uncertain and imprecise. The Robust Optimization and Probabilistic Analysis of Robustness (ROPAR) algorithm is a well-suited tool for dealing with uncertainty. Still, the failure to consider multiple uncertainties and multi-objective robustness hinders the application of the ROPAR algorithm to practical problems. This paper proposes a robust optimization and robustness probabilistic analysis method that considers numerous uncertainties and multi-objective robustness for robust water resources allocation under uncertainty. The copula function is introduced for analyzing the probabilities of different scenarios. The robustness with respect to the two objective functions is analyzed separately, and the Pareto frontier of robustness is generated. The relationship between the robustness with respect to the two objective functions is used to evaluate water resources management strategies. Use of the method is illustrated in a case study of water resources allocation in the Huaihe River basin. The results demonstrate that the method opens a possibility for water managers to make more informed uncertainty-aware decisions. ...
Journal article (2024) - Shaokun He, Yi Bo Wang, Dimitri Solomatine, Xiao Li
Long-term water resource management involving multipurpose coordination requires robust decision-making in water infrastructure cases to cope with various types of uncertainties. Traditional robust optimization methods generally do not explicitly propagate input or parametric uncertainties into estimates of the robustness of solutions, which limits their ability to address uncertainty comprehensively across solution spaces. In this study, we introduce an explicit robust decision-making framework that blends multiobjective search, probabilistic analysis of robustness, and diagnostic verification tools to identify robust optimal solutions to external uncertainty. The proposed framework is illustrated on four diverse robustness formulations, which capture a wide variety of stakeholder attitudes from highly risk-averse to risk-neutral, for the primary operating objectives (hydropower production, water diversion, and hydrological alteration degree) in China's Hanjiang cascade reservoir system. By analyzing the Pareto front propagated from inflow uncertainty, it is found that optimal robust policies with a significantly higher degree of hydrological alteration are preferred in most formulations to achieve relatively lower joint uncertainty of hydropower and water diversion. These policies also yield sufficiently stable model performance in the case of an out-of-sample streamflow set during diagnostic verification. Furthermore, a comparative analysis of four different formulations suggests that a composite normalized robustness indicator (NRI) developed in this study to integrate various robustness metrics can achieve an effective balance for all considered objectives. These findings highlight the benefits of explicit robust optimization for managing hydrological uncertainties in multipurpose cascade reservoirs. ...
Journal article (2024) - Ana M. Paez-Trujillo, J. Sebastian Hernandez-Suarez, Leonardo Alfonso, Beatriz Hernandez, Shreedhar Maskey, Dimitri Solomatine
While drought impacts are widespread across the globe, climate change projections indicate more frequent and severe droughts. This underscores the pressing need to increase resistance and resilience to drought. The strategic application of Preventive Drought Management Measures (PDMMs) is a suitable avenue to reduce the likelihood of drought and ameliorate associated damages. In this study, we use an optimisation approach with a multicriteria decision-making method to allocate PDMMs for reducing the severity of agricultural and hydrological droughts. The results indicate that implementing PDMMs can reduce the severity of agricultural and hydrological droughts, and the obtained management scenarios (solutions) highlight the utility of multi-objective optimisation for PDMMs planning. However, examined management scenarios also illustrate the trade-off between managing agricultural and hydrological droughts. PDMMs can alleviate the severity of agricultural droughts while producing opposite effects for hydrological droughts (or vice versa). Furthermore, the impact of PDMMs displays temporal and spatial variabilities. For instance, PDMMs implementation within a specific subbasin may mitigate the severity of one type of drought in a given month yet exacerbate drought conditions in preceding or subsequent months. In the case of hydrological droughts, the PDMMs may intensify streamflow deficits in the intervened subbasins while alleviating the hydrological drought severity downstream (or vice versa). These complexities emphasise a customised implementation of PDMMs, considering the basin characteristics (e.g., rainfall distribution over the year, soil properties, land use, and topography) and the quantification of PDMMs' effect on the severity of each type of drought. ...
Journal article (2024) - Renata Sadrtdinova, Gerald Augusto Corzo Perez, Dimitri P. Solomatine
Kazakhstan is recently experiencing an increase in drought trends. However, low-capacity probabilistic drought forecasts and poor dissemination have led to a drought crisis in 2021 that resulted in the loss of thousands of livestock. To improve drought forecasting accuracy, this study applies Machine Learning and Deep Learning (ML and DL) algorithms to capture the sequences of drought events using a non-contiguous drought analysis (NCDA). Precipitation, 2-m temperature, runoff, solar radiation, relative humidity, and evaporation were collected from the ERA5 database as input variables. Combinations of inputs were used to build ML models, including seven classifiers (Logistic, K-NN, Kernel SVM, Decision Tree, Random Forest, XGBoost, and GRU). The output events were defined by standardized precipitation index (SPI) and SPEI indicators as binary classes. Weekly time series from 1991 to 2021 for each cell were used to forecast a lead time from 1 week to 6 months. GRU provided 97–99% accuracy in more volatile regions while Random Forest and XGBoost showed 94–99% accuracy at a lead time of 6 months. The accuracy evaluation was based on the confusion matrix and F1 score to analyze the stage change capture. This study demonstrates the effectiveness of using ML and DL algorithms for drought forecasting, with potential applications for other regions. ...

Cluster Size Filter and Drought Indicator Threshold Optimization

Book chapter (2024) - Vitali Diaz, Gerald A. Corzo Perez, Henny A.J. Van Lanen, Dimitri P. Solomatine
In its three-dimensional (3-D) characterization, drought is an event whose spatial extent changes over time. Each drought event has an onset and end time, a location, a magnitude, and a spatial trajectory. These characteristics help to analyze and describe how drought develops in space and time (i.e., drought dynamics). Methodologies for 3-D characterization of drought include a 3-D clustering technique to extract the drought events from the hydrometeorological data. The application of the clustering method yields small artifact droughts. These small clusters are removed from the analysis with the use of a cluster size filter. However, according to the literature, the filter parameters are usually set arbitrarily, so this study concentrated on a method to calculate the optimal cluster size filter for the 3-D characterization of drought. The effect of different drought indicator thresholds to calculate drought is also analyzed. The approach was tested in South America with data from the Latin American Flood and Drought Monitor for 1950–2017. Analysis of the spatial trajectories and characteristics of the most extreme droughts is also included. Calculated droughts are compared with information reported at a country scale and a reasonably good match is found. ...
Journal article (2023) - Ana Paez Trujillo, Jeffer Cañon, Beatriz Hernandez, Gerald Corzo, Dimitri Solomatine
The typical drivers of drought events are lower than normal precipitation and/or higher than normal evaporation. The region's characteristics may enhance or alleviate the severity of these events. Evaluating the combined effect of the multiple factors influencing droughts requires innovative approaches. This study applies hydrological modelling and a machine learning tool to assess the relationship between hydroclimatic characteristics and the severity of agricultural and hydrological droughts. The Soil Water Assessment Tool (SWAT) is used for hydrological modelling. Model outputs, soil moisture and streamflow, are used to calculate two drought indices, namely the Soil Moisture Deficit Index and the Standardized Streamflow Index. Then, drought indices are utilised to identify the agricultural and hydrological drought events during the analysis period, and the index categories are employed to describe their severity. Finally, the multivariate regression tree technique is applied to assess the relationship between hydroclimatic characteristics and the severity of agricultural and hydrological droughts.

Our research indicates that multiple parameters influence the severity of agricultural and hydrological droughts in the Cesar River basin. The upper part of the river valley is very susceptible to agricultural and hydrological drought. Precipitation shortfalls and high potential evapotranspiration drive severe agricultural drought, whereas limited precipitation influences severe hydrological drought. In the middle part of the river, inadequate rainfall partitioning and an unbalanced water cycle that favours water loss through evapotranspiration and limits percolation cause severe agricultural and hydrological drought conditions. Finally, droughts are moderate in the basin's southern part (Zapatosa marsh and the Serranía del Perijá foothills). Moderate sensitivity to agricultural and hydrological droughts is related to the capacity of the subbasins to retain water, which lowers evapotranspiration losses and promotes percolation. Results show that the presented methodology, combining hydrological modelling and a machine learning tool, provides valuable information about the interplay between the hydroclimatic factors that influence drought severity in the Cesar River basin. ...
Book chapter (2023) - Santiago Duarte, Gerald A. Corzo Perez, Germán Santos, Dimitri P. Solomatine
Human behavior and decision making are dynamically influenced by digital media, making digital news a vast data source of events and points of view. Meanwhile, artificial intelligence now has advanced capacity for natural language processing (NLP) through tools such as sentiment analysis and topic identification. This research uses machine learning algorithms to prove the hypothesis that it is possible to find correlations between news information and water resource problems. The 207 water bodies in the Magdalena River Basin, Colombia, were analyzed alongside 19,490 news articles published between 2016 and 2020 in 42 online newspapers. A platform for the visualization of spatiotemporal information was developed as a proof of concept. There has been a noticeable increase in digital news about the Magdalena River Basin in recent years and some correlation with extreme events, but the comparison between the sentiments and measured physical variables was inconclusive. This is likely due to the difficulty in filtering the topic from obtained news, as well as in precisely identifying the spatial location and temporal range of events. In the future, NLP techniques to analyze news articles could be used to provide additional information for decision making at the basin level. ...
Journal article (2023) - Paul Muñoz, Gerald Corzo, Dimitri Solomatine, Jan Feyen, Rolando Célleri
Extreme peak runoff forecasting is still a challenge in hydrology. In fact, the use of traditional physically-based models is limited by the lack of sufficient data and the complexity of the inner hydrological processes. Here, we employ a Machine Learning technique, the Random Forest (RF) together with a combination of Feature Engineering (FE) strategies for adding physical knowledge to RF models and improving their forecasting performances. The FE strategies include precipitation-event classification according to hydrometeorological criteria and separation of flows into baseflow and directflow. We used ∼ 3.5 years of hourly precipitation information retrieved from two near-real-time satellite precipitation databases (PERSIANN-CCS and IMERG-ER), and runoff data at the outlet of a 3391-km2 basin located in the tropical Andes of Ecuador. The developed models obtained Nash-Sutcliffe efficiencies varying from 0.86 to 0.59 for lead times between 1 and 6 h. The best performances were obtained for peak runoffs triggered by short-extension precipitation events (<50 km2) where infiltration- or saturation-excess runoff responses are well learned by the RF models. Conversely, the forecasting difficulty is associated with extensive precipitation events. For such conditions, a deeper characterization of the biophysical characteristics of the basin is encouraged for capturing the dynamic of directflow across multiple runoff responses. All in all, the potential to employ near-real-time satellite precipitation and the use of FE strategies for improving RF forecasting provides hydrologists with new tools for real-time runoff forecasting in remote or complex regions. ...
Journal article (2023) - Clara Maria Corzo, Leonardo Alfonso, Gerald Corzo, Dimitri Solomatine
Water utilities are urged to decrease their real water losses, not only to reduce costs but also to assure long-term sustainability. Hardware- and software-based techniques have been broadly used to locate leaks; within the latter, previous works that have used data-driven models mostly focused on single leaks. This paper presents a methodology to locate multiple leaks in water distribution networks employing pressure residuals. It consists of two phases: one is to produce training data for the data-driven model and cluster the nodes based on their leak-flow-rate-independent signatures using an adapted hierarchical agglomerative algorithm; the second is to locate the leaks using a top-down approach. To identify the leaking clusters and nodes, we employed a custom-built k-nearest neighbor (k-NN) algorithm that compares the test instances with the generated training data. This instance-to-instance comparison requires substantial computational resources for classification, which was overcome by the use of high-performance computing. The methodology was applied to a real network located in a European town, comprising 144 nodes and a total length of pipes of 24 km. Although its multiple inlets add redundancy to the network increasing the challenge of leak location, the method proved to obtain acceptable results to guide the field pinpointing activities. Nearly 70% of the areas determined by the clusters were identified with an accuracy of over 90% for leak flows above 3.0 L/s, and the leaking nodes were accurately detected over 50% of the time for leak flows above 4.0 L/s. ...
Book chapter (2023) - Mostafa Farrag, Gerald A.Corzo Perez, Dimitri P. Solomatine
Conceptual hydrological models imply a simplification of the complexity of the hydrological system; however, they lack the flexibility in reproducing a wide range of the catchment responses. Usually, a trade-off is done to sacrifice the accuracy of a specific aspect of the system behavior in favor of the accuracy of other aspects. This study evaluates the benefit of using a modular approach, “The fuzzy committee model” of building specialized models to reproduce specific responses of the catchment. We also assess the applicability of using predicted runoff from specialized models to form a fuzzy committee model. In this paper, weighting schemes with power parameter values are investigated. A thorough study is conducted on the relation between the fuzzy committee variables (the membership functions and the weighting schemes), and their effect on the model performance. Furthermore, the Fuzzy committee concept is applied on a conceptual distributed model with two cases, the first with lumped catchment parameters and the latter with distributed parameters. A comparison between different combinations of the fuzzy committee variables showed the superiority of all Fuzzy Committee models over single models. Fuzzy committee of distributed models performed well, especially in capturing the highest peak in the calibration data set; however, it needs further study of the effect of model parameterization on the model performance and uncertainty. ...
Journal article (2023) - Emmanouil A. Varouchakis, Dimitri Solomatine, Gerald A. Corzo Perez, Seifeddine Jomaa, George P. Karatzas
Successful modelling of the groundwater level variations in hydrogeological systems in complex formations considerably depends on spatial and temporal data availability and knowledge of the boundary conditions. Geostatistics plays an important role in model-related data analysis and preparation, but has specific limitations when the aquifer system is inhomogeneous. This study combines geostatistics with machine learning approaches to solve problems in complex aquifer systems. Herein, the emphasis is given to cases where the available dataset is large and randomly distributed in the different aquifer types of the hydrogeological system. Self-Organizing Maps can be applied to identify locally similar input data, to substitute the usually uncertain correlation length of the variogram model that estimates the correlated neighborhood, and then by means of Transgaussian Kriging to estimate the bias corrected spatial distribution of groundwater level. The proposed methodology was tested on a large dataset of groundwater level data in a complex hydrogeological area. The obtained results have shown a significant improvement compared to the ones obtained by classical geostatistical approaches. ...