R. Van de Plas
Please Note
28 records found
1
The Hilbert-Huang transform is able to decompose the complex nonlinear wave structure caused by ship movement. Previous research studies that used conventional frequency analysis methods are insufficient to extract local time-frequency characteristics of the waves due to their nonlinearity and non-stationary behavior. The Hilbert-Huang transform combines Empirical Mode Decomposition (EMD) with Hilbert spectral analysis to extract local time-frequency information of the wave.
EMD can be extended to Ensemble EMD, where the method is improved using white noise. This extension solves the mode mixing problem commonly encountered in EMD, resulting in more physically meaningful Intrinsic Mode Functions (IMFs). Once the decomposition is complete, the extracted wave components—corresponding to the primary and secondary wave structures—are identified and grouped based on their frequency characteristics.
These extracted wave components are used as inputs for regression models such as regression tree algorithms and neural network regression models. By training these models on the decomposed wave data, it becomes possible to describe and predict critical wave characteristics as a function of relevant input parameters, such as ship speed, ship-gauge distance, and navigation channel symmetry. ...
The Hilbert-Huang transform is able to decompose the complex nonlinear wave structure caused by ship movement. Previous research studies that used conventional frequency analysis methods are insufficient to extract local time-frequency characteristics of the waves due to their nonlinearity and non-stationary behavior. The Hilbert-Huang transform combines Empirical Mode Decomposition (EMD) with Hilbert spectral analysis to extract local time-frequency information of the wave.
EMD can be extended to Ensemble EMD, where the method is improved using white noise. This extension solves the mode mixing problem commonly encountered in EMD, resulting in more physically meaningful Intrinsic Mode Functions (IMFs). Once the decomposition is complete, the extracted wave components—corresponding to the primary and secondary wave structures—are identified and grouped based on their frequency characteristics.
These extracted wave components are used as inputs for regression models such as regression tree algorithms and neural network regression models. By training these models on the decomposed wave data, it becomes possible to describe and predict critical wave characteristics as a function of relevant input parameters, such as ship speed, ship-gauge distance, and navigation channel symmetry.
We investigate computational methods for two specific imaging instruments: (1) a time-of-flight imaging mass spectrometer, employed in biochemical research to visualize molecular distributions across very small organic tissues, and (2) a mid-infrared imager, utilized in astronomical research to study very large protostars, temperate exoplanets, and objects within our solar system. Despite their considerable promise in acquiring detailed molecular maps and critical astronomical insights, respectively, the practical analysis and interpretation of their image sets face substantial obstacles, namely their dimensionality and the effect of noise. Addressing these challenges may involve drawing on existing computational and storage capacity and harnessing any available prior information or problem-specific structure. Fortunately, analytical solutions to the obstacles across imaging instruments often bear a resemblance to each other, as we distil them to abstract mathematical models and eventually formulate those problems as optimization problems.
The computational methods we are interested in are so-called low-rankmethods, they can simultaneously provide insight in data structure (analysis), as well as reduce the dimensionality of the data and denoise it.... ...
We investigate computational methods for two specific imaging instruments: (1) a time-of-flight imaging mass spectrometer, employed in biochemical research to visualize molecular distributions across very small organic tissues, and (2) a mid-infrared imager, utilized in astronomical research to study very large protostars, temperate exoplanets, and objects within our solar system. Despite their considerable promise in acquiring detailed molecular maps and critical astronomical insights, respectively, the practical analysis and interpretation of their image sets face substantial obstacles, namely their dimensionality and the effect of noise. Addressing these challenges may involve drawing on existing computational and storage capacity and harnessing any available prior information or problem-specific structure. Fortunately, analytical solutions to the obstacles across imaging instruments often bear a resemblance to each other, as we distil them to abstract mathematical models and eventually formulate those problems as optimization problems.
The computational methods we are interested in are so-called low-rankmethods, they can simultaneously provide insight in data structure (analysis), as well as reduce the dimensionality of the data and denoise it....
Chemical Imaging Methods for Cultural Heritage
Advanced Data Acquisition and Processing
Firstly, chapter 1 introduces the basic concepts of cultural heritage and the field of cultural heritage science. It also provides an overview of some of the scientific methods used for the study of cultural heritage objects, with a focus on imaging methods, and particularly imaging spectroscopy methods. The two primary methods considered in this work, macro X-ray fluorescence spectroscopy (MA-XRF) and reflectance imaging spectroscopy (RIS), are described in greater detail and an overview of the state-of-the-art in equipment and data processing methods is provided.
Beginning with the issue of data processing, chapter 2 discusses a novel method for the use of short-wave infrared (SWIR) (1000–2500nm) RIS for semi-quantitative analysis of historical paintings. The method consists of the isolation and deconvolution of characteristic absorption features of target pigments. The concept is proven on two pigments, lead white and blue verditer. The method is tested on a set of specially prepared paint samples as well as a 16th century painting. The method is compared to MA-XRF on its ability to selectively map pigments as well as its ability to provide quantitative information on the pigment concentrations. A novel data visualization method, able to visualize chemical relevance of individual pixels whilst also highlighting larger spatial patterns, is also presented.
Chapter 3 continues on the topic of data processing, presenting a novel approach for the analysis of damaged historical manuscripts through the use of visible and nearinfrared (VNIR) (400–1000 nm) RIS combined with supervised machine learning methods and machine learning explainability methods. The approach uses manually labeled VNIR RIS data acquired from historical manuscripts to train an XGBoost classification model, followed by the application of Shapley additive explanations (SHAP) to analyse the behaviour of the classification model. The calculated SHAP values are then used to calculate a SHAP-weighted intensity map (SWIM), which is found to improve the legibility of the analysed manuscripts. An adaptive colour scheme is also proposed as a method of easing the evaluation by paleographists of the resulting images. The approach is tested on two texts, the Leiden Riddle, a 9-10th century Northumbrian text, and the 1669 First Set of the Fundamental Constitutions of Carolina, the earliest text attributed to political philosopher John Locke.
Shifting to the issue of data acquisition, chapter 4 documents the design and testing of two MA-XRF scanners. These scanners, the big lead box (BLB) Mark I and Mark II, are designed to be flexiblemeasurement platforms that allow for the testing of alternate scanning strategies and multi-modal acquisitions. Highlighted are the safety features implemented into the design of the scanners with the goal of minimizing potential radiation exposure to users and bystanders.
Following the development of the MA-XRF scanners, chapter 5 discusses two methods for the acceleration of MA-XRF measurements through the use of smart scanning strategies. The first method, the Fast Autonomous Scanning Toolkit (FAST), uses machine learning models to dynamically select which pixels to scan whilst trying to reduce the estimated distortion between the reconstructed data set and the underlying ground truth. The second method, the Chopp algorithm, uses a double scan approach, where a first fast initial scan is used to estimate the optimal scan time per pixel for a second scan with a specified total scan time. The two methods are evaluated on their scan times and the achieved data quality compared to each other and a traditional raster scan.
Lastly, chapter 6 provides some concluding remarks on the themes covered during this work and highlights potential future developments. ...
Firstly, chapter 1 introduces the basic concepts of cultural heritage and the field of cultural heritage science. It also provides an overview of some of the scientific methods used for the study of cultural heritage objects, with a focus on imaging methods, and particularly imaging spectroscopy methods. The two primary methods considered in this work, macro X-ray fluorescence spectroscopy (MA-XRF) and reflectance imaging spectroscopy (RIS), are described in greater detail and an overview of the state-of-the-art in equipment and data processing methods is provided.
Beginning with the issue of data processing, chapter 2 discusses a novel method for the use of short-wave infrared (SWIR) (1000–2500nm) RIS for semi-quantitative analysis of historical paintings. The method consists of the isolation and deconvolution of characteristic absorption features of target pigments. The concept is proven on two pigments, lead white and blue verditer. The method is tested on a set of specially prepared paint samples as well as a 16th century painting. The method is compared to MA-XRF on its ability to selectively map pigments as well as its ability to provide quantitative information on the pigment concentrations. A novel data visualization method, able to visualize chemical relevance of individual pixels whilst also highlighting larger spatial patterns, is also presented.
Chapter 3 continues on the topic of data processing, presenting a novel approach for the analysis of damaged historical manuscripts through the use of visible and nearinfrared (VNIR) (400–1000 nm) RIS combined with supervised machine learning methods and machine learning explainability methods. The approach uses manually labeled VNIR RIS data acquired from historical manuscripts to train an XGBoost classification model, followed by the application of Shapley additive explanations (SHAP) to analyse the behaviour of the classification model. The calculated SHAP values are then used to calculate a SHAP-weighted intensity map (SWIM), which is found to improve the legibility of the analysed manuscripts. An adaptive colour scheme is also proposed as a method of easing the evaluation by paleographists of the resulting images. The approach is tested on two texts, the Leiden Riddle, a 9-10th century Northumbrian text, and the 1669 First Set of the Fundamental Constitutions of Carolina, the earliest text attributed to political philosopher John Locke.
Shifting to the issue of data acquisition, chapter 4 documents the design and testing of two MA-XRF scanners. These scanners, the big lead box (BLB) Mark I and Mark II, are designed to be flexiblemeasurement platforms that allow for the testing of alternate scanning strategies and multi-modal acquisitions. Highlighted are the safety features implemented into the design of the scanners with the goal of minimizing potential radiation exposure to users and bystanders.
Following the development of the MA-XRF scanners, chapter 5 discusses two methods for the acceleration of MA-XRF measurements through the use of smart scanning strategies. The first method, the Fast Autonomous Scanning Toolkit (FAST), uses machine learning models to dynamically select which pixels to scan whilst trying to reduce the estimated distortion between the reconstructed data set and the underlying ground truth. The second method, the Chopp algorithm, uses a double scan approach, where a first fast initial scan is used to estimate the optimal scan time per pixel for a second scan with a specified total scan time. The two methods are evaluated on their scan times and the achieved data quality compared to each other and a traditional raster scan.
Lastly, chapter 6 provides some concluding remarks on the themes covered during this work and highlights potential future developments.
Compressed Sensing for Sparse Channel Estimation in Next Generation Radios
Addressing Hardware Impairments
Randomized Nonnegative Matrix Factorization for Imaging Mass Spectrometry
Exploration of Nonnegative Bases for Entry-wise Nonnegative Projection
In this thesis, the author demonstrates how Randomized NMF inherits its two-stage framework for computation from Randomized SVD. The main bottleneck between Randomized NMF replicating the success of its predecessor is the lower-dimensional basis used to compress the data matrix. Although the basis has been conventionally effective for Randomized SVD with provable guarantees, the lack of entry-wise nonnegativity constraint requires existing methods to compromise the projection quality and interpretability and rely on the relaxation of nonnegativity constraints within the update rules. Therefore, this work motivates and explores the use of alternative bases that directly satisfy entry-wise nonnegativity. ...
In this thesis, the author demonstrates how Randomized NMF inherits its two-stage framework for computation from Randomized SVD. The main bottleneck between Randomized NMF replicating the success of its predecessor is the lower-dimensional basis used to compress the data matrix. Although the basis has been conventionally effective for Randomized SVD with provable guarantees, the lack of entry-wise nonnegativity constraint requires existing methods to compromise the projection quality and interpretability and rely on the relaxation of nonnegativity constraints within the update rules. Therefore, this work motivates and explores the use of alternative bases that directly satisfy entry-wise nonnegativity.
Unlike many clustering algorithms that have assumptions about the geometry of the clusters, spectral clustering constructs a similarity graph and performs eigendecomposition on the Laplacian matrix to reveal non-convex clusters. This allows one to find clusters of arbitrary shapes, which can result in new or improved segmentation being discovered in IMS data. Furthermore, two recent studies allow for the potential argument that spectral clustering might be optimal for IMS data.
Despite these advantages, spectral clustering faces implementation barriers primarily due to its computational complexity and memory constraints. The limited applications of spectral clustering on IMS data, can predominantly be attributed to these limitations.
This thesis investigates the feasibility of spectral clustering for analyzing high-dimensional imaging mass spectrometry data, with a focus on performance under noise, computational scalability, and maintaining biological segmentation. To assess the performance, internal and external validation metrics are used as well as a comparison with variations of k-means clustering. Additionally, a memory constrained algorithm was developed to address the scalability issue induced by the memory complexity.
The results highlight that spectral clustering outperforms k-means, when both methods are utilizing the cosine metric, in scenarios of increased noise on a synthetic dataset. Upon application on a real world subset of an IMS dataset of a mouse pup containing the brain, the results between k-means and spectral clustering were highly comparable. When applied on the complete dataset with a memory constrained version of spectral clustering, the results were less promising due to its dependence on initial seeding, where k-means obtained better or similar clustering results with lower time and memory complexity. ...
Unlike many clustering algorithms that have assumptions about the geometry of the clusters, spectral clustering constructs a similarity graph and performs eigendecomposition on the Laplacian matrix to reveal non-convex clusters. This allows one to find clusters of arbitrary shapes, which can result in new or improved segmentation being discovered in IMS data. Furthermore, two recent studies allow for the potential argument that spectral clustering might be optimal for IMS data.
Despite these advantages, spectral clustering faces implementation barriers primarily due to its computational complexity and memory constraints. The limited applications of spectral clustering on IMS data, can predominantly be attributed to these limitations.
This thesis investigates the feasibility of spectral clustering for analyzing high-dimensional imaging mass spectrometry data, with a focus on performance under noise, computational scalability, and maintaining biological segmentation. To assess the performance, internal and external validation metrics are used as well as a comparison with variations of k-means clustering. Additionally, a memory constrained algorithm was developed to address the scalability issue induced by the memory complexity.
The results highlight that spectral clustering outperforms k-means, when both methods are utilizing the cosine metric, in scenarios of increased noise on a synthetic dataset. Upon application on a real world subset of an IMS dataset of a mouse pup containing the brain, the results between k-means and spectral clustering were highly comparable. When applied on the complete dataset with a memory constrained version of spectral clustering, the results were less promising due to its dependence on initial seeding, where k-means obtained better or similar clustering results with lower time and memory complexity.
The Kerr nonlinear effects add phase shifts to the signal, which are dependent on its instantaneous power. These phase distortions occur simultaneously with the dispersion effect of the fiber, which spreads signal pulses in time. The interplay is complicated and makes compensation of distortions challenging. The nonlinear Fourier transform(NFT), which offers immunity from the distortions of the Kerr effect, received great interest in recent years. The lossless nonlinear Schrödinger equation (NLSE), which models signal propagation in an ideal lossless optical fiber, belongs to a class of nonlinear partial differential equations known as integrable equations. These integrable equations can be solved exactly by NFT. Similar to the Fourier transform that translates a linear dispersive propagation in the time domain into phase delays in the signal spectrum, the NFT translates the nonlinear evolution of the signal governed by the lossless NLSE into trivial multiplications in the nonlinear Fourier spectrum of the signal. The NFT is exact for lossless fiber channels. In the presence of loss, the integrability property is violated. In lossy propagation, signal power reduces as it propagates. This in turn reduces the strength of the nonlinear effects along the length of the fiber. As practical fibers are lossy, the path-average approximation is often used to apply NFT on lossy fiber channels. In this approximation, the variation in the Kerr nonlinear effects due to the reduction in signal power is accounted as the variations in the Kerr-nonlinearity parameter of the fiber. Then, by approximating the varying Kerr-nonlinearity parameter with its average value over a span, a lossless fiber model is obtained. This approximation has errors associated with it which sacrifices the performance. We developed a NFT-based transmission system that is exact even in the presence of fiber loss. The proposed design eliminates errors due to loss, thus improving performance over the design that uses path-average approximation… ...
The Kerr nonlinear effects add phase shifts to the signal, which are dependent on its instantaneous power. These phase distortions occur simultaneously with the dispersion effect of the fiber, which spreads signal pulses in time. The interplay is complicated and makes compensation of distortions challenging. The nonlinear Fourier transform(NFT), which offers immunity from the distortions of the Kerr effect, received great interest in recent years. The lossless nonlinear Schrödinger equation (NLSE), which models signal propagation in an ideal lossless optical fiber, belongs to a class of nonlinear partial differential equations known as integrable equations. These integrable equations can be solved exactly by NFT. Similar to the Fourier transform that translates a linear dispersive propagation in the time domain into phase delays in the signal spectrum, the NFT translates the nonlinear evolution of the signal governed by the lossless NLSE into trivial multiplications in the nonlinear Fourier spectrum of the signal. The NFT is exact for lossless fiber channels. In the presence of loss, the integrability property is violated. In lossy propagation, signal power reduces as it propagates. This in turn reduces the strength of the nonlinear effects along the length of the fiber. As practical fibers are lossy, the path-average approximation is often used to apply NFT on lossy fiber channels. In this approximation, the variation in the Kerr nonlinear effects due to the reduction in signal power is accounted as the variations in the Kerr-nonlinearity parameter of the fiber. Then, by approximating the varying Kerr-nonlinearity parameter with its average value over a span, a lossless fiber model is obtained. This approximation has errors associated with it which sacrifices the performance. We developed a NFT-based transmission system that is exact even in the presence of fiber loss. The proposed design eliminates errors due to loss, thus improving performance over the design that uses path-average approximation…
Deep NMF can decompose IMS data into multiple structures, helping to uncover complex biochemical interactions. This thesis explores the application of deep NMF to IMS data, deriving update algorithms for both the Frobenius norm and generalized Kullback-Leibler (gKL) divergence, and testing these methods on MALDI-TOF IMS data. Results show that the deep multiplicative update rule (MUR) outperforms non-deep methods in terms of interpretability and surpasses multi-layer structures in both reconstruction error and interpretability. Computational challenges remain in tuning algorithm parameters and achieving convergence in deep Alternating Direction Method of Multipliers (ADMM) NMF.
...
Deep NMF can decompose IMS data into multiple structures, helping to uncover complex biochemical interactions. This thesis explores the application of deep NMF to IMS data, deriving update algorithms for both the Frobenius norm and generalized Kullback-Leibler (gKL) divergence, and testing these methods on MALDI-TOF IMS data. Results show that the deep multiplicative update rule (MUR) outperforms non-deep methods in terms of interpretability and surpasses multi-layer structures in both reconstruction error and interpretability. Computational challenges remain in tuning algorithm parameters and achieving convergence in deep Alternating Direction Method of Multipliers (ADMM) NMF.
In this thesis, we suggest a 1-D CNN architecture that extracts local features along the spectral dimension. The aim is to investigate if CNNs improve the classification accuracy compared to other classic Machine Learning (ML) methods such as linear models. Furthermore, we explore Neural Networks (NNs) that employ the novel Sharpened Cosine Similarity (SCS) as a feature extraction method, opposed to convolution. We call those networks SCS-NN in correspondence to the Convolutional-NN (CNN). To evaluate these methods, we implement our pipeline for various IMS datasets, with different characteristics and classification tasks, using several performance metrics such as balanced accuracy and F1 score.
Moreover, we provide a detailed description of the methodology pipeline used for the CNN architecture design. The suggested methodology is the Tree-structured Parzen Estimator (TPE) algorithm, a Bayesian optimization technique for automated architecture selection. By implementing TPE, we manage to explore and exploit efficiently a complex and large hyperparameter configuration space and automatically select optimal hyperparameters (such as number of convolutional layers, kernel size, strides, learning rates etc.). This automated approach reduces time consumption, errors, and the need for specialized knowledge in biology and biochemistry that would be associated with manual design. In addition to developing a pipeline for designing, training and evaluating a CNN for IMS data classification, we also apply a model agnostic interpretation methodology based on SHapley Additive exPlanations (SHAP) and provide SHAP score maps that visualize the importance of features in the spatial dimension of the IMS datacube.
In this thesis, we present and analyse the automated selection of 1-D CNN architectures for IMS data classification based on the TPE algorithm. Furthermore, we investigate a novel alternative to convolution, SCS, and evaluate its strengths and weaknesses in IMS data classification. The experimental results show that the TPE-generated CNN architectures outperform all the other applied classifiers. Finally, our interpretation of the CNN models reveals that accuracy performance alone might not be a sufficient criterion to trust the model's output. ...
In this thesis, we suggest a 1-D CNN architecture that extracts local features along the spectral dimension. The aim is to investigate if CNNs improve the classification accuracy compared to other classic Machine Learning (ML) methods such as linear models. Furthermore, we explore Neural Networks (NNs) that employ the novel Sharpened Cosine Similarity (SCS) as a feature extraction method, opposed to convolution. We call those networks SCS-NN in correspondence to the Convolutional-NN (CNN). To evaluate these methods, we implement our pipeline for various IMS datasets, with different characteristics and classification tasks, using several performance metrics such as balanced accuracy and F1 score.
Moreover, we provide a detailed description of the methodology pipeline used for the CNN architecture design. The suggested methodology is the Tree-structured Parzen Estimator (TPE) algorithm, a Bayesian optimization technique for automated architecture selection. By implementing TPE, we manage to explore and exploit efficiently a complex and large hyperparameter configuration space and automatically select optimal hyperparameters (such as number of convolutional layers, kernel size, strides, learning rates etc.). This automated approach reduces time consumption, errors, and the need for specialized knowledge in biology and biochemistry that would be associated with manual design. In addition to developing a pipeline for designing, training and evaluating a CNN for IMS data classification, we also apply a model agnostic interpretation methodology based on SHapley Additive exPlanations (SHAP) and provide SHAP score maps that visualize the importance of features in the spatial dimension of the IMS datacube.
In this thesis, we present and analyse the automated selection of 1-D CNN architectures for IMS data classification based on the TPE algorithm. Furthermore, we investigate a novel alternative to convolution, SCS, and evaluate its strengths and weaknesses in IMS data classification. The experimental results show that the TPE-generated CNN architectures outperform all the other applied classifiers. Finally, our interpretation of the CNN models reveals that accuracy performance alone might not be a sufficient criterion to trust the model's output.
Accelerating MA-XRF Data Acquisition by Exploiting Local Spatial and Spectral Relations within a Hyperspectral Datacube
An Approach through Wavelet Denoising
Data-driven Parameter Optimization and Spatial Awareness for Uniform Manifold Approximation and Projection (UMAP) of IMS data sets
A step towards parameter-free dimensionality reduction
In this thesis, we introduce UMAPLUS, an extension of UMAP that takes not only spectral, but also spatial information into account. Our experiments show that UMAPLUS is capable of extracting the structure of a synthetic IMS data set better than UMAP. Besides UMAP’s standard spectral and random initializations, we introduce a new spatial initialization that provides more intuitive insight into the LD embedding relative to the spatial image domain. Standard UMAP entails several parameters that influence consistency, reliability, and quality of its LD embedding. Thus far, the influence of these parameters on IMS data sets has been barely investigated, and previous studies have consistently applied the default settings of UMAP. In this thesis, a Data-Driven UMAP (DD-UMAP) is constructed, which includes optimization of UMAP parameters in a data-driven manner. Several distance metrics are investigated, with cosine similarity providing the most robust results. A naive 1-D optimization procedure is compared to a multivariate Bayesian optimization approach, capable of optimizing multiple parameters simultaneously. Using an evaluation function partly based on the cost function of UMAP and the addition of spatial information, DD-UMAP is able to estimate and utilize an optimized input parameter set for UMAP in an unsupervised manner and on unlabeled IMS data sets. This is demonstrated on generated synthetic data, and two real world IMS data sets; a full mouse pup and a murine kidney. ...
In this thesis, we introduce UMAPLUS, an extension of UMAP that takes not only spectral, but also spatial information into account. Our experiments show that UMAPLUS is capable of extracting the structure of a synthetic IMS data set better than UMAP. Besides UMAP’s standard spectral and random initializations, we introduce a new spatial initialization that provides more intuitive insight into the LD embedding relative to the spatial image domain. Standard UMAP entails several parameters that influence consistency, reliability, and quality of its LD embedding. Thus far, the influence of these parameters on IMS data sets has been barely investigated, and previous studies have consistently applied the default settings of UMAP. In this thesis, a Data-Driven UMAP (DD-UMAP) is constructed, which includes optimization of UMAP parameters in a data-driven manner. Several distance metrics are investigated, with cosine similarity providing the most robust results. A naive 1-D optimization procedure is compared to a multivariate Bayesian optimization approach, capable of optimizing multiple parameters simultaneously. Using an evaluation function partly based on the cost function of UMAP and the addition of spatial information, DD-UMAP is able to estimate and utilize an optimized input parameter set for UMAP in an unsupervised manner and on unlabeled IMS data sets. This is demonstrated on generated synthetic data, and two real world IMS data sets; a full mouse pup and a murine kidney.
On the Atoms of Robustness
Robust Matrix Decomposition for Spectral Imaging
Data-Driven Soft Discriminant Maps
Class-aware Linear Feature Extraction in Imaging Mass Spectrometry
Since some medical decisions are made based on imaging, separating the biological signal from noise is of significant importance (e.g. accelerates decision-making, reducing the chance of misdiagnosis).
Some non-biological variations that span a wide range of imaging modalities include e.g. viewport stitching artifacts, slice-to-slice interference, aliasing, and Gibbs-phenomena.
From a signal processing perspective, many of these can be modeled as quasiperiodic patterns.
Thus, removal of quasiperiodic patterns while preserving the underlying medical information is the main focus of this thesis.
Although in modern instruments, many forms of non-biological variation can be attenuated to be invisible to the naked eye, machine learning algorithms which are often used for classification of disease and segmentation of biological samples may be susceptible to even minor variations and noise patterns.
Development of entirely data-driven, unsupervised denoising techniques can potentially increase the effectiveness and reliability of such algorithms.
Furthermore, under certain image transformations, such as different color spaces, the Fourier and wavelet transform, and factorizations, such as principal component analysis and non-negative matrix factorization, as well as combinations of these, non-biological patterns can get amplified and become so prominent that much of the underlying biological information is concealed.
Removing these quasiperiodic patterns using current state-of-the-art algorithms still requires manual parameter tuning and prior expert knowledge, which is an impractical and possibly unnecessary expectation towards healthcare professionals.
The goal of this M.Sc. thesis is to develop an automated, data-driven framework, that is able to reliably identify, quantify, and eliminate quasiperiodic patterns within the images while retaining as much biological information as possible.
In this framework, named Quasiperiodic Image Denoising (QID), two novel algorithms are implemented, both operating in the Fourier domain.
One algorithm is based on robust principal component analysis (QID-RPCA) and the other uses the normalized median of absolute differences (QID-MADN).
The methods used to achieve unsupervised, data-driven denoising are described in detail.
This includes the use of histogram equalization for radial binning, an automated, sparsity-based approach to choosing the optimal aggregation level, and noise component attenuation based on radial frequency patterns.
The methodology is demonstrated through three case studies.
First, a synthetic dataset is used to compare the performance of the novel algorithms to the current state-of-the-art solutions.
Second, the performance is evaluated on two real-world datasets processed using a number of methods, e.g. factorization and different color spaces. One of these datasets is based on a microscopy image of a transversal section of a mouse brain and the other one is based on a microscopy image of a coronal section of a rat kidney.
Finally, a real-world, raw dataset is denoised consisting of a set of high-resolution fluorescent microscopy images of a human kidney.
Results indicate that the novel algorithms have higher denoising performance than previous approaches in the literature with notable improvements achieved for low-frequency corruptions. ...
Since some medical decisions are made based on imaging, separating the biological signal from noise is of significant importance (e.g. accelerates decision-making, reducing the chance of misdiagnosis).
Some non-biological variations that span a wide range of imaging modalities include e.g. viewport stitching artifacts, slice-to-slice interference, aliasing, and Gibbs-phenomena.
From a signal processing perspective, many of these can be modeled as quasiperiodic patterns.
Thus, removal of quasiperiodic patterns while preserving the underlying medical information is the main focus of this thesis.
Although in modern instruments, many forms of non-biological variation can be attenuated to be invisible to the naked eye, machine learning algorithms which are often used for classification of disease and segmentation of biological samples may be susceptible to even minor variations and noise patterns.
Development of entirely data-driven, unsupervised denoising techniques can potentially increase the effectiveness and reliability of such algorithms.
Furthermore, under certain image transformations, such as different color spaces, the Fourier and wavelet transform, and factorizations, such as principal component analysis and non-negative matrix factorization, as well as combinations of these, non-biological patterns can get amplified and become so prominent that much of the underlying biological information is concealed.
Removing these quasiperiodic patterns using current state-of-the-art algorithms still requires manual parameter tuning and prior expert knowledge, which is an impractical and possibly unnecessary expectation towards healthcare professionals.
The goal of this M.Sc. thesis is to develop an automated, data-driven framework, that is able to reliably identify, quantify, and eliminate quasiperiodic patterns within the images while retaining as much biological information as possible.
In this framework, named Quasiperiodic Image Denoising (QID), two novel algorithms are implemented, both operating in the Fourier domain.
One algorithm is based on robust principal component analysis (QID-RPCA) and the other uses the normalized median of absolute differences (QID-MADN).
The methods used to achieve unsupervised, data-driven denoising are described in detail.
This includes the use of histogram equalization for radial binning, an automated, sparsity-based approach to choosing the optimal aggregation level, and noise component attenuation based on radial frequency patterns.
The methodology is demonstrated through three case studies.
First, a synthetic dataset is used to compare the performance of the novel algorithms to the current state-of-the-art solutions.
Second, the performance is evaluated on two real-world datasets processed using a number of methods, e.g. factorization and different color spaces. One of these datasets is based on a microscopy image of a transversal section of a mouse brain and the other one is based on a microscopy image of a coronal section of a rat kidney.
Finally, a real-world, raw dataset is denoised consisting of a set of high-resolution fluorescent microscopy images of a human kidney.
Results indicate that the novel algorithms have higher denoising performance than previous approaches in the literature with notable improvements achieved for low-frequency corruptions.