J. Sun
Please Note
24 records found
1
GFMs for Biodiversity
Adapting Geospatial Foundation Models for Biodiversity Monitoring
Mutational signatures in the general population
Population-scale mutational signature analysis of blood-derived genomes from the UK Biobank
Personalized event prediction applied to (medical) time-series
Using per-patient Isolation Forest anomaly scores as features
Three personalization methods were implemented: z-score normalization, Empirical Cumulative Distribution Function (ECDF) percentile scoring, and Isolation Forest anomaly detection. Each personalization method constructed patient-specific features by comparing current windows against the patient’s
own historical distribution, also called their baseline. These features were evaluated using static models (XGBoost and Logistic Regression) and a dynamic model (Temporal Convolutional Network) using ECG and PPG waveform data from the MIMIC-III database to predict sepsis and mortality hours in advance.
The results show that adding personalized features consistently improves predictive performance when using a sliding baseline over using global features alone. The AUROC when predicting sepsis improved from 0.72 to 0.82 when adding z-score normalized features, and mortality prediction AUROC improved from 0.77 to 0.82 when adding the Isolation Forest anomaly score. The optimal baseline length is dependent on the task. Using all available prior data is best for predicting sepsis, while a baseline constructed from the previous 4-8 hours is optimal for mortality prediction. A sliding baseline, which uses the most recent history for each prediction window, consistently outperforms a fixed baseline placed at the start of the stay.
The Isolation Forest anomaly score, which was computed individually per patient, was the most important feature for both sepsis and mortality prediction when combining all personalized feature sets with global features. Unlike other personalized features, the anomaly score also improved Logistic Regression performance, suggesting a more linear relationship with mortality risk.
Models trained on PPG data show similar improvement to ECG-based models, and approximately the same optimal baseline lengths are found across both signal types, for both sepsis and mortality prediction. This suggests that the optimal baseline is event-specific rather than just dataset-specific. The strong performance of PPG data is particularly promising because wearable devices could provide
a healthy pre-admission personal baseline, potentially mitigating the cold start problem and improving
performance further.
These results demonstrate that incorporating personal historical baselines into clinical prediction models is a practical and effective approach to improving early warning systems in the ICU. These methods are likely broadly applicable to any time-series domain where individual patterns carry predictive value. ...
Three personalization methods were implemented: z-score normalization, Empirical Cumulative Distribution Function (ECDF) percentile scoring, and Isolation Forest anomaly detection. Each personalization method constructed patient-specific features by comparing current windows against the patient’s
own historical distribution, also called their baseline. These features were evaluated using static models (XGBoost and Logistic Regression) and a dynamic model (Temporal Convolutional Network) using ECG and PPG waveform data from the MIMIC-III database to predict sepsis and mortality hours in advance.
The results show that adding personalized features consistently improves predictive performance when using a sliding baseline over using global features alone. The AUROC when predicting sepsis improved from 0.72 to 0.82 when adding z-score normalized features, and mortality prediction AUROC improved from 0.77 to 0.82 when adding the Isolation Forest anomaly score. The optimal baseline length is dependent on the task. Using all available prior data is best for predicting sepsis, while a baseline constructed from the previous 4-8 hours is optimal for mortality prediction. A sliding baseline, which uses the most recent history for each prediction window, consistently outperforms a fixed baseline placed at the start of the stay.
The Isolation Forest anomaly score, which was computed individually per patient, was the most important feature for both sepsis and mortality prediction when combining all personalized feature sets with global features. Unlike other personalized features, the anomaly score also improved Logistic Regression performance, suggesting a more linear relationship with mortality risk.
Models trained on PPG data show similar improvement to ECG-based models, and approximately the same optimal baseline lengths are found across both signal types, for both sepsis and mortality prediction. This suggests that the optimal baseline is event-specific rather than just dataset-specific. The strong performance of PPG data is particularly promising because wearable devices could provide
a healthy pre-admission personal baseline, potentially mitigating the cold start problem and improving
performance further.
These results demonstrate that incorporating personal historical baselines into clinical prediction models is a practical and effective approach to improving early warning systems in the ICU. These methods are likely broadly applicable to any time-series domain where individual patterns carry predictive value.
Data augmentation for Sparse Graph Traversals
Exploring data augmentation options to enhance deep learning model performance
Despite limited improvements, this study establishes a foundation for future research in graph-based trajectory augmentation. Integrating richer trip-level features, such as dynamic environmental conditions or behavioral data, with structural augmentation could lead to more effective training data and improved model generalization. ...
Despite limited improvements, this study establishes a foundation for future research in graph-based trajectory augmentation. Integrating richer trip-level features, such as dynamic environmental conditions or behavioral data, with structural augmentation could lead to more effective training data and improved model generalization.
Data augmentation for graph based data
Improving representation of cycling trips with varying speed conditions using data augmentation
How well can machine learning tools for humanitarian forecasting be used in predicting the consequences of forced displacement?
Humanitarian forecasting for displacement: a survey
Machine learning for humanitarian forecasting: A Survey
Assessing the trustworthiness and real-world feasibility of machine learning models for conflict forecasting
Impact-based humanitarian forecasting using machine learning for floods
A literature survey
Field-Based Predictive Growth Modeling of Broccoli (Brassica oleracea var. italica) within Precision Agriculture
Understanding Field Growth Dynamics for Data-Driven Agriculture
This thesis investigates how field-based measurements can be used to model broccoli growth at the level of individual plants. The work addresses four problems: converting raw field video into plant-level growth curves, investigating which environmental parameters best describe the growth, determining whether cumulative temperature or thermal time better represents broccoli development, and evaluating how different growth models capture head diameter growth. A preliminary study using an external dataset creates the methodological foundation by analysing broccoli growth and benchmarking classical and neural models. A field study in a Dutch production environment expands this work through the development of a data-processing pipeline that includes head detection, plant identification, tracking, diameter estimation, and integration with local weather data. Across both studies, thermal time provides a more biologically meaningful predictor of development than cumulative temperature. Model comparison shows that classical parametric models capture general developmental trends, while a multi-layer perceptron achieves the highest predictive accuracy, with a mean absolute error of 0.583~cm, when multiple environmental variables are included. The results demonstrate that field-based predictive modeling can support precision-agriculture applications such as harvest planning and selective mechanized harvesting. ...
This thesis investigates how field-based measurements can be used to model broccoli growth at the level of individual plants. The work addresses four problems: converting raw field video into plant-level growth curves, investigating which environmental parameters best describe the growth, determining whether cumulative temperature or thermal time better represents broccoli development, and evaluating how different growth models capture head diameter growth. A preliminary study using an external dataset creates the methodological foundation by analysing broccoli growth and benchmarking classical and neural models. A field study in a Dutch production environment expands this work through the development of a data-processing pipeline that includes head detection, plant identification, tracking, diameter estimation, and integration with local weather data. Across both studies, thermal time provides a more biologically meaningful predictor of development than cumulative temperature. Model comparison shows that classical parametric models capture general developmental trends, while a multi-layer perceptron achieves the highest predictive accuracy, with a mean absolute error of 0.583~cm, when multiple environmental variables are included. The results demonstrate that field-based predictive modeling can support precision-agriculture applications such as harvest planning and selective mechanized harvesting.
Towards Smarter Greenhouses: Combining Physics and Machine Learning
Evaluating the Impact & Opportunities of Physics-Informed Machine Learning on the Task of Greenhouse Humidity Prediction
Edge-aware Bilateral Filtering
Reducing across-edge blurring for the bilateral filter
On-Mesh Bilateral Filtering
Bridging the Gap Between Texture and Object Space
Predictable blur behaviour for the bilateral filter
Researching a method for linear behaviour between the blurriness and spatial filter size of the bilateral filter
Detecting Patterns in Train Position Data of Trains in Shunting Yards
Analysis of Arrival Time Distributions and Delays
This research focuses on extracting and analyzing the arrival times of trains at shunting yards using a dataset consisting of GPS data. It conducts two algorithms to cluster the given data for each train unit within and across days to identify the same train across different days. Distributions and heatmaps of the arrival times and delays are created based on the identified train series. They are analyzed to identify patterns in train arrival times and delays across different months. ...
This research focuses on extracting and analyzing the arrival times of trains at shunting yards using a dataset consisting of GPS data. It conducts two algorithms to cluster the given data for each train unit within and across days to identify the same train across different days. Distributions and heatmaps of the arrival times and delays are created based on the identified train series. They are analyzed to identify patterns in train arrival times and delays across different months.