RG

R. Ghorbani

info

Please Note

8 records found

Towards Reliable Time Series Anomaly Detection

Doctoral thesis (2025) - R. Ghorbani, M.J.T. Reinders, D.M.J. Tax
The integration of wearable technology into healthcare is revolutionizing health monitoring by enabling continuous tracking of vital metrics like heart rate and blood sugar. Devices such as smartwatches and glucose monitors empower proactive interventions, reducing hospital visits and personalizing care. For instance, wearables can detect irregular heart rhythms for early cardiovascular disease detection or assist individuals with diabetes in managing glucose levels. These advancements are enabled by technologies like photoplethysmography (PPG), a non-invasive method for real-time monitoring of physiological signals. Continuous monitoring generates time series data that captures dynamic health fluctuations over time. This data allows for identifying irregularities and deviations that isolated measurements might miss. Detecting anomalies, such as abrupt changes in heart rate or prolonged abnormal patterns, is essential for timely interventions in managing chronic conditions like hypertension and cardiovascular diseases.

However, the analysis of time series data introduces challenges. For instance, label scarcity arises because labeling health anomalies requires expert input, which is often infeasible for large datasets. Inter-subject variability becomes a concern as physiological patterns differ significantly across individuals, complicating model generalization. Furthermore, temporal dependencies in time series data add complexity, as observations are sequentially related and anomalies may not manifest as isolated points but as patterns or sequences deviating from normal behavior. Detecting subtle anomalies, minor deviations that accumulate over time but may signal early-stage conditions, becomes particularly challenging due to their resemblance to normal temporal variations and noise in the data. For example, a gradual change in heart rate variability might indicate the onset of an irregular rhythm but could easily blend into inherent variability if not carefully analyzed. Moreover, evaluation metrics for time series data are insufficient, failing to capture the temporal complexities of real-world applications. Consequently, conventional metrics can misrepresent model performance, leading to unreliable or misleading assessments.

This thesis addresses these challenges by advancing time series anomaly detection through innovative methodologies. A key focus is on addressing the limitations of existing evaluation metrics by introducing new evaluation metrics to better capture temporal complexities, ensuring reliable and meaningful performance assessments. Beyond evaluation, this thesis is guided by several core principles to address the challenges inherent in time series data analysis. Central to this is the use of unsupervised representation learning to tackle label scarcity and variability, enabling robust feature extraction from unlabeled data while maintaining generalizability. Finally, the thesis develops strategies for increasing sensitivity to subtle anomalies, providing effective solutions for identifying small yet significant deviations in complex datasets. Together, these contributions present a comprehensive framework for improving anomaly detection systems across diverse applications, bridging theoretical advancements with practical, real-world needs. ...
Anomaly detection in time series data is crucial across various domains. The scarcity of labeled data for such tasks has increased the attention towards unsupervised learning methods. These approaches, often relying solely on reconstruction error, typically fail to detect subtle anomalies in complex datasets. To address this, we introduce RESTAD, an adaptation of the Transformer model by incorporating a layer of Radial Basis Function (RBF) neurons within its architecture. This layer fits a non-parametric density in the latent representation, such that a high RBF output indicates similarity with predominantly normal training data. RESTAD integrates the RBF similarity scores with the reconstruction errors to increase sensitivity to anomalies. Our empirical evaluations demonstrate that RESTAD outperforms various established baselines across multiple benchmark datasets. ...
Photoplethysmography (PPG) signals, typically acquired from wearable devices, hold significant potential for continuous fitness-health monitoring. In particular, heart conditions that manifest in rare and subtle deviating heart patterns may be interesting. However, robust and reliable anomaly detection within these data remains a challenge due to the scarcity of labeled data and high inter-subject variability. This paper introduces a two-stage framework leveraging representation learning and personalization to improve anomaly detection performance in PPG data. The proposed framework first employs representation learning to transform the original PPG signals into a more discriminative and compact representation. We then apply three different unsupervised anomaly detection methods for movement detection and biometric identification. We validate our approach using two different datasets in both generalized and personalized scenarios. Our results demonstrate significant improvements: for movement detection, in the generalized scenario, AUCs improved from barely 0.5 to above 0.9 with representation learning. Importantly, inter-subject variability was substantially reduced, from around 0.4 to below 0.1. In the personalized scenario, AUCs became close to 1.0, with variability further reduced to below 0.05, indicating the effectiveness of both representation learning and personalization for anomaly detection in PPG data. Similar enhancements were observed in biometric identification, emphasizing how our approach can minimize inter-subject variability and enhance PPG-based health monitoring systems. ...

Proximity-Aware Time Series Anomaly Evaluation

Evaluating anomaly detection algorithms in time series data is critical as inaccuracies can lead to flawed decision-making in various domains where real-time analytics and data-driven strategies are essential. Traditional performance metrics assume iid data and fail to capture the complex temporal dynamics and specific characteristics of time series anomalies, such as early and delayed detections. We introduce Proximity-Aware Time series anomaly Evaluation (PATE), a novel evaluation metric that incorporates the temporal relationship between prediction and anomaly intervals. PATE uses proximity-based weighting considering buffer zones around anomaly intervals, enabling a more detailed and informed assessment of a detection. Using these weights, PATE computes a weighted version of the area under the Precision and Recall curve. Our experiments with synthetic and real-world datasets show the superiority of PATE in providing more sensible and accurate evaluations than other evaluation metrics. We also tested several state-of-the-art anomaly detectors across various benchmark datasets using the PATE evaluation scheme. The results show that a common metric like Point-Adjusted F1 Score fails to characterize the detection performances well, and that PATE is able to provide a more fair model comparison. By introducing PATE, we redefine the understanding of model efficacy that steers future studies toward developing more effective and accurate detection models. ...
With the progress of sensor technology in wearables, the collection and analysis of PPG signals are gaining more interest. Using Machine Learning, the cardiac rhythm corresponding to PPG signals can be used to predict different tasks such as activity recognition, sleep stage detection, or more general health status. However, supervised learning is often limited by the amount of available labeled data, which is typically expensive to obtain. To address this problem, we propose a Self-Supervised Learning (SSL) method with a pretext task of signal reconstruction to learn an informative generalized PPG representation. The performance of the proposed SSL framework is compared with two fully supervised baselines. The results show that in a very limited label data setting (10 samples per class or less), using SSL is beneficial, and a simple classifier trained on SSL-learned representations outperforms fully supervised deep neural networks. However, the results reveal that the SSL-learned representations are too focused on encoding the subjects. Unfortunately, there is high inter-subject variability in the SSL-learned representations, which makes working with this data more challenging when labeled data is scarce. The high inter-subject variability suggests that there is still room for improvements in learning representations. In general, the results suggest that SSL may pave the way for the broader use of machine learning models on PPG data in label-scarce regimes. ...
Conference paper (2022) - M. Moradi, R. Ghorbani, Stefano Sfarra, D.M.J. Tax, D. Zarouchas
Assessment of cultural heritage assets is now extremely important all around the world. Non-destructive inspection is essential for preserving the integrity of the artworks while avoiding the loss of any precious materials that make it up. The use of Infrared Thermography (IRT) is an interesting concept since surface and subsurface faults can be discovered by utilizing the 3D diffusion inside the object caused by external heat. The primary goal of this research is to detect defects in artworks, which is one of the most important tasks in the restoration of mural paintings. To this end, a spatiotemporal deep neural network (STDNN) is utilized for defect identification in a mock-up reproducing an artwork, taking into account both the temporal and spatial perspectives of step-heating (SH) thermography. Finally, the outcomes are compared to those of other conventional algorithms. ...
Journal article (2022) - M. Moradi, R. Ghorbani, Stefano Sfarra, D.M.J. Tax, D. Zarouchas
Assessment of cultural heritage assets is now extremely important all around the world. Non-destructive inspection is essential for preserving the integrity of artworks while avoiding the loss of any precious materials that make them up. The use of Infrared Thermography is an interesting concept since surface and subsurface faults can be discovered by utilizing the 3D diffusion inside the object caused by external heat. The primary goal of this research is to detect defects in artworks, which is one of the most important tasks in the restoration of mural paintings. To this end, machine learning and deep learning techniques are effective tools that should be employed properly in accordance with the experiment’s nature and the collected data. Considering both the temporal and spatial perspectives of step-heating thermography, a spatiotemporal deep neural network is developed for defect identification in a mock-up reproducing an artwork. The results are then compared with those of other conventional algorithms, demonstrating that the proposed approach outperforms the others. ...
Journal article (2020) - Ramin Ghorbani, Rouzbeh Ghousi, Ahmad Makui, Alireza Atashi
Due to the development of biomedical equipment and healthcare level, especially in the Intensive Care Unit (ICU), a considerable amount of data has been collected for analysis. Mortality prediction in the ICUs is considered as one of the most important topics in the healthcare data analysis section. A precise prediction of the mortality risk for patients in ICU could provide us with valuable information about patients' lives and reduce costs at the earliest possible stage. This paper aims to introduce a new hybrid predictive model using the Genetic Algorithm as a feature selection method and a new ensemble classifier based on the combination of Stacking and Boosting ensemble methods to create an early mortality prediction model on a highly imbalanced dataset. The SVM-SMOTE method is used to solve the imbalanced data problem. This paper compares the new model with various machine learning models to validate the efficiency of the introduced model. The achieved results using the shuffle 5-fold cross-validation and random hold-out methods indicate that the new hybrid model has the best performance among other classifiers. Additionally, the Friedman test is applied as a statistical significance test to examine the differences between classifiers. The results of the statistical analysis prove that the proposed model is more effective than other classifiers. Furthermore, the proposed model is compared to APACHE and SAPS scoring systems and is benchmarked against state-of-the-art predictive models applied to the MIMIC dataset for experimental validation and achieved promising results as it outperformed the state-of-the-art models. ...