RV

R. Van de Plas

info

Please Note

28 records found

High-dimensional spectral imaging generates data cubes in which each pixel contains a full spectrum rather than a limited set of channels. Techniques such as hyperspectral imaging, tomography, and mass spectrometry produce rich spatial and spectral information, enabling detailed analysis of composition and structure across applications in science and engineering. Examples include imaging mass spectrometry (IMS) for molecular analysis of biological tissues and cosmological experiments such as TIME, which use spectral data to study large-scale structures in the universe. While these methods offer significant analytical power, they also introduce challenges related to the storage, processing, and interpretation of large, high-dimensional datasets. ...
Recent observations show that ship-induced waves in navigation channels (i.e. shallow water) damage river bank revetments. This has opened new research into understanding the physical behavior of ship-induced shallow water waves. This study is conducted using data gathered by the German Federal Institute for Hydraulic Engineering (BAW). The data is obtained from laboratory measurements and analyzed using the Hilbert-Huang transform, followed by complex regression methods, to enable prediction of the wave components.

The Hilbert-Huang transform is able to decompose the complex nonlinear wave structure caused by ship movement. Previous research studies that used conventional frequency analysis methods are insufficient to extract local time-frequency characteristics of the waves due to their nonlinearity and non-stationary behavior. The Hilbert-Huang transform combines Empirical Mode Decomposition (EMD) with Hilbert spectral analysis to extract local time-frequency information of the wave.

EMD can be extended to Ensemble EMD, where the method is improved using white noise. This extension solves the mode mixing problem commonly encountered in EMD, resulting in more physically meaningful Intrinsic Mode Functions (IMFs). Once the decomposition is complete, the extracted wave components—corresponding to the primary and secondary wave structures—are identified and grouped based on their frequency characteristics.

These extracted wave components are used as inputs for regression models such as regression tree algorithms and neural network regression models. By training these models on the decomposed wave data, it becomes possible to describe and predict critical wave characteristics as a function of relevant input parameters, such as ship speed, ship-gauge distance, and navigation channel symmetry. ...
Doctoral thesis (2026) - R.A.R. Moens, R. Van de Plas, B. De Schutter
A common strategy to address new scientific challenges consists of abstracting the underlying problem, recasting it to an existing problem formulation and applying an established methodology. In this dissertation, we offer a variation on this familiar academic theme. The setting we will focus on is primarily found within image-acquiring instruments, characterized by producing vast quantities of data, from several hundreds up to more than half a million images per experiment. The challenges that we address throughout this work will mainly consist of (a) reducing dimensionality and (b) denoising, which have a direct and significant impact on the analysis and thus interpretation of these extensive image sets.
We investigate computational methods for two specific imaging instruments: (1) a time-of-flight imaging mass spectrometer, employed in biochemical research to visualize molecular distributions across very small organic tissues, and (2) a mid-infrared imager, utilized in astronomical research to study very large protostars, temperate exoplanets, and objects within our solar system. Despite their considerable promise in acquiring detailed molecular maps and critical astronomical insights, respectively, the practical analysis and interpretation of their image sets face substantial obstacles, namely their dimensionality and the effect of noise. Addressing these challenges may involve drawing on existing computational and storage capacity and harnessing any available prior information or problem-specific structure. Fortunately, analytical solutions to the obstacles across imaging instruments often bear a resemblance to each other, as we distil them to abstract mathematical models and eventually formulate those problems as optimization problems.
The computational methods we are interested in are so-called low-rankmethods, they can simultaneously provide insight in data structure (analysis), as well as reduce the dimensionality of the data and denoise it.... ...

Advanced Data Acquisition and Processing

Imaging spectroscopy methods are becoming increasingly relevant in the field of cultural heritage science. The datacubes output by these methods represent some of the most significant challenges related to their application, namely howto make sense of complex multi-dimensional datasets. As the selection of imaging spectroscopy methods and the complexity of the resulting datasets continue to grow, the time required to conduct all these measurements and process all of the data also increases. This research focuses on addressing both the problem of extended data acquisition times and the challenge of processing the complex datacubes.

Firstly, chapter 1 introduces the basic concepts of cultural heritage and the field of cultural heritage science. It also provides an overview of some of the scientific methods used for the study of cultural heritage objects, with a focus on imaging methods, and particularly imaging spectroscopy methods. The two primary methods considered in this work, macro X-ray fluorescence spectroscopy (MA-XRF) and reflectance imaging spectroscopy (RIS), are described in greater detail and an overview of the state-of-the-art in equipment and data processing methods is provided.

Beginning with the issue of data processing, chapter 2 discusses a novel method for the use of short-wave infrared (SWIR) (1000–2500nm) RIS for semi-quantitative analysis of historical paintings. The method consists of the isolation and deconvolution of characteristic absorption features of target pigments. The concept is proven on two pigments, lead white and blue verditer. The method is tested on a set of specially prepared paint samples as well as a 16th century painting. The method is compared to MA-XRF on its ability to selectively map pigments as well as its ability to provide quantitative information on the pigment concentrations. A novel data visualization method, able to visualize chemical relevance of individual pixels whilst also highlighting larger spatial patterns, is also presented.

Chapter 3 continues on the topic of data processing, presenting a novel approach for the analysis of damaged historical manuscripts through the use of visible and nearinfrared (VNIR) (400–1000 nm) RIS combined with supervised machine learning methods and machine learning explainability methods. The approach uses manually labeled VNIR RIS data acquired from historical manuscripts to train an XGBoost classification model, followed by the application of Shapley additive explanations (SHAP) to analyse the behaviour of the classification model. The calculated SHAP values are then used to calculate a SHAP-weighted intensity map (SWIM), which is found to improve the legibility of the analysed manuscripts. An adaptive colour scheme is also proposed as a method of easing the evaluation by paleographists of the resulting images. The approach is tested on two texts, the Leiden Riddle, a 9-10th century Northumbrian text, and the 1669 First Set of the Fundamental Constitutions of Carolina, the earliest text attributed to political philosopher John Locke.

Shifting to the issue of data acquisition, chapter 4 documents the design and testing of two MA-XRF scanners. These scanners, the big lead box (BLB) Mark I and Mark II, are designed to be flexiblemeasurement platforms that allow for the testing of alternate scanning strategies and multi-modal acquisitions. Highlighted are the safety features implemented into the design of the scanners with the goal of minimizing potential radiation exposure to users and bystanders.

Following the development of the MA-XRF scanners, chapter 5 discusses two methods for the acceleration of MA-XRF measurements through the use of smart scanning strategies. The first method, the Fast Autonomous Scanning Toolkit (FAST), uses machine learning models to dynamically select which pixels to scan whilst trying to reduce the estimated distortion between the reconstructed data set and the underlying ground truth. The second method, the Chopp algorithm, uses a double scan approach, where a first fast initial scan is used to estimate the optimal scan time per pixel for a second scan with a specified total scan time. The two methods are evaluated on their scan times and the achieved data quality compared to each other and a traditional raster scan.

Lastly, chapter 6 provides some concluding remarks on the themes covered during this work and highlights potential future developments. ...
Doctoral thesis (2026) - H. Masoumi, R. Van de Plas, N.J. Myers
Millimeter wave (mmWave) bands, currently used in 5G, offer significant spectrum to enable Gbps data rates for emerging data-hungry applications. At mmWave and higher bands, however, high scattering and atmospheric attenuation result in low received power. To address this issue, base stations employ large antenna arrays and periodically configure these arrays by learning the propagation environment, called the wireless channel. The use of such large arrays, however, makes the learning process, i.e., wireless channel estimation, extremely challenging. This is because, on the one hand, the training overhead of classical channel estimation methods increases significantly with the array dimensions. On the other hand, compressed sensing (CS)based methods for fast channel estimation suffer from a poor signal-to-noise ratio (SNR) in the received measurements. Moreover, radio frequency (RF) impairments, which are much more severe in mmWave systems than in low-frequency systems, further complicate channel estimation. For example, low-resolution RF phase shifters, phase noise at the oscillator, and in-phase/quadrature (IQ) mismatch at the down-conversion are among the practically significant impairments.... ...

Exploration of Nonnegative Bases for Entry-wise Nonnegative Projection

Master thesis (2025) - A.A. Mukhamedov, Raf Van de Plas, R.A.R. Moens
Imaging Mass Spectrometry (IMS) data calls for computational methods to perform factorizations efficiently that facilitate interpretable components. Despite the widespread use of Nonnegative Matrix Factorization (NMF) for IMS-related dimensionality reduction, conventional NMF algorithms are not up to par to scale for extremely large datasets. To remedy this problem, this thesis explores Randomized Nonnegative Matrix Factorization as an efficient alternative to compute IMS data on whole-body mouse pup tissue data.

In this thesis, the author demonstrates how Randomized NMF inherits its two-stage framework for computation from Randomized SVD. The main bottleneck between Randomized NMF replicating the success of its predecessor is the lower-dimensional basis used to compress the data matrix. Although the basis has been conventionally effective for Randomized SVD with provable guarantees, the lack of entry-wise nonnegativity constraint requires existing methods to compromise the projection quality and interpretability and rely on the relaxation of nonnegativity constraints within the update rules. Therefore, this work motivates and explores the use of alternative bases that directly satisfy entry-wise nonnegativity. ...
Imaging Mass Spectrometry (IMS) collects spatial and chemical information of a sample, generating high-dimensional datasets that present challenges in exploratory data analysis due to their substantial size. Spectral clustering is a promising unsupervised learning approach for IMS applications, employing graph-based strategies to identify patterns without assumptions about cluster geometry.

Unlike many clustering algorithms that have assumptions about the geometry of the clusters, spectral clustering constructs a similarity graph and performs eigendecomposition on the Laplacian matrix to reveal non-convex clusters. This allows one to find clusters of arbitrary shapes, which can result in new or improved segmentation being discovered in IMS data. Furthermore, two recent studies allow for the potential argument that spectral clustering might be optimal for IMS data.

Despite these advantages, spectral clustering faces implementation barriers primarily due to its computational complexity and memory constraints. The limited applications of spectral clustering on IMS data, can predominantly be attributed to these limitations.

This thesis investigates the feasibility of spectral clustering for analyzing high-dimensional imaging mass spectrometry data, with a focus on performance under noise, computational scalability, and maintaining biological segmentation. To assess the performance, internal and external validation metrics are used as well as a comparison with variations of k-means clustering. Additionally, a memory constrained algorithm was developed to address the scalability issue induced by the memory complexity.

The results highlight that spectral clustering outperforms k-means, when both methods are utilizing the cosine metric, in scenarios of increased noise on a synthetic dataset. Upon application on a real world subset of an IMS dataset of a mouse pup containing the brain, the results between k-means and spectral clustering were highly comparable. When applied on the complete dataset with a memory constrained version of spectral clustering, the results were less promising due to its dependence on initial seeding, where k-means obtained better or similar clustering results with lower time and memory complexity. ...
Doctoral thesis (2025) - V. Bajaj, M.H.G. Verhaegen, S. Wahls, R. Van de Plas
The global internet protocol (IP) traffic is rising exponentially due to the increased number of bandwidth-intensive services like video-on-demand and cloud computing. To meet the demands of growing traffic, the data rates of fiber optic communication systems (FOCSs) need to be increased. In this regard, digital signal processing (DSP), which already plays a powerful role in the modern FOCSs, is being explored. Increasing the net data rate of FOCSs requires compensation for nonlinear impairments that can arise from the Kerr effect during the propagation of signal through fiber as well as from the non-ideal responses of transceiver hardware components. Furthermore, cutbacks in net data rates due to training overheads, i.e., the non-information carrying part of transmitted data consumed by DSP algorithms; need to be reduced. In this dissertation, we propose novel DSP approaches to address these problems of great practical interest.

The Kerr nonlinear effects add phase shifts to the signal, which are dependent on its instantaneous power. These phase distortions occur simultaneously with the dispersion effect of the fiber, which spreads signal pulses in time. The interplay is complicated and makes compensation of distortions challenging. The nonlinear Fourier transform(NFT), which offers immunity from the distortions of the Kerr effect, received great interest in recent years. The lossless nonlinear Schrödinger equation (NLSE), which models signal propagation in an ideal lossless optical fiber, belongs to a class of nonlinear partial differential equations known as integrable equations. These integrable equations can be solved exactly by NFT. Similar to the Fourier transform that translates a linear dispersive propagation in the time domain into phase delays in the signal spectrum, the NFT translates the nonlinear evolution of the signal governed by the lossless NLSE into trivial multiplications in the nonlinear Fourier spectrum of the signal. The NFT is exact for lossless fiber channels. In the presence of loss, the integrability property is violated. In lossy propagation, signal power reduces as it propagates. This in turn reduces the strength of the nonlinear effects along the length of the fiber. As practical fibers are lossy, the path-average approximation is often used to apply NFT on lossy fiber channels. In this approximation, the variation in the Kerr nonlinear effects due to the reduction in signal power is accounted as the variations in the Kerr-nonlinearity parameter of the fiber. Then, by approximating the varying Kerr-nonlinearity parameter with its average value over a span, a lossless fiber model is obtained. This approximation has errors associated with it which sacrifices the performance. We developed a NFT-based transmission system that is exact even in the presence of fiber loss. The proposed design eliminates errors due to loss, thus improving performance over the design that uses path-average approximation… ...
Master thesis (2024) - A.C.C. Molenaars, R. Van de Plas, R.A.R. Moens
In medical research, Imaging Mass Spectrometry (IMS) is a powerful tool that facilitates the spatial mapping of biomolecules in tissue samples, contributing to the identification of disease biomarkers and the analysis of drug effects. While offering valuable insights, IMS data is often large and complex, presenting challenges in data interpretation. Traditional techniques such as principal component analysis (PCA) and non-negative matrix factorization (NMF) have been employed to address these issues, but they fall short due to their inability to fully capture the complex, mixed nature of IMS data. Deep non-negative matrix factorization (deep NMF) offers a promising solution by employing a multi-layer architecture, allowing for dimensionality reduction while improving interpretability.

Deep NMF can decompose IMS data into multiple structures, helping to uncover complex biochemical interactions. This thesis explores the application of deep NMF to IMS data, deriving update algorithms for both the Frobenius norm and generalized Kullback-Leibler (gKL) divergence, and testing these methods on MALDI-TOF IMS data. Results show that the deep multiplicative update rule (MUR) outperforms non-deep methods in terms of interpretability and surpasses multi-layer structures in both reconstruction error and interpretability. Computational challenges remain in tuning algorithm parameters and achieving convergence in deep Alternating Direction Method of Multipliers (ADMM) NMF.
...
This report investigates the use of Generative Adversarial Nets (GANs) specifically for over- sampling Imaging Mass Spectrometry spectra. IMS is a technique used to measure the spatial distribution of molecules, which is valuable in fields like oncology and biomarker discovery. GANs, on the other hand, are a class of machine learning frameworks where two neural net- works, the generator, and the discriminator, are trained simultaneously through adversarial processes. The generator creates synthetic data, while the discriminator tries to distinguish between real and synthetic data. GANs-based oversampling aims to increase classifier performance by adding data to classes that are underrepresented in the original data. Synthetic oversampling is especially relevant in IMS data as the measuring technique is destructive, making acquiring more real samples impossible. GANs have been shown to outperform other oversampling techniques such as SMOTE on various datasets. Applying GANs directly to the dataset proved unsuccessful in this oversampling task. Different possible causes of the limited performance of the GANs are studied leading to improved experiment results using spectra reduced in dimension and the Wasserstein GANs with gradient penalty. Even though with these changes to the experiment the GANs appear to generate more realistic data, using this data for oversampling does not increase overall classifier performance. Rather, it steers the classifier to overfitting towards the minority classes. This report demonstrates that applying the designed GANs for oversampling minority classes on this dataset does increase classifier performance. However, it is shown that GANs can be trained on IMS data and that GANs might be of use for applications with IMS data besides oversampling. ...
Master thesis (2024) - G. Sinha, R. Van de Plas, L.G. Migas
The main aim of the project was to develop a novel algorithm that enable two-dimensional feature detection in an extremely sparse environment of Ion Mobility Imaging Mass Spectrometry (IM-IMS) measurements. For this, 2D Wavelet Transform Maxima is proposed. This led to construction of wavelet chains (ridges) in the wavelet transform space. Current thresholding methods for these chains were based on global thresholding of wavelet coefficients. Therefore, a thresholding method known as effective length-based thresholding is introduced. Here, length-based thresholding is modified to take the local nature of noise into account. The design parameters of the algorithm were evaluated using a synthetic IM-IMS data sample, and the performance of the algorithm was quantified using F-Score. The parameters that were studied were (i) width of the wavelet function (along mobility dimension), (ii) effective length, (iii) penalty factor used for calculation of local noise. Fair performance was obtained by choosing effective length equal to 4 and width equal to 1% of mobility dimension. There were a significant number of false positives still being detected along dominant mobilograms . The performance of the algorithm was compared with an existing feature detection algorithm on a real-world IM-IMS data sample and was found to be better in terms of number of features being detected. ...
Convolutional Neural Networks (CNNs) have emerged primarily from research focusing on image classification tasks and as a result, most of the well-motivated design choices found in literature are relevant to computer vision applications. CNNs' application on Imaging Mass Spectrometry (IMS) data is quite recent and involves new challenges, such as taking into account their unique structure (e.g. both spatial and spectral dimensions).

In this thesis, we suggest a 1-D CNN architecture that extracts local features along the spectral dimension. The aim is to investigate if CNNs improve the classification accuracy compared to other classic Machine Learning (ML) methods such as linear models. Furthermore, we explore Neural Networks (NNs) that employ the novel Sharpened Cosine Similarity (SCS) as a feature extraction method, opposed to convolution. We call those networks SCS-NN in correspondence to the Convolutional-NN (CNN). To evaluate these methods, we implement our pipeline for various IMS datasets, with different characteristics and classification tasks, using several performance metrics such as balanced accuracy and F1 score.

Moreover, we provide a detailed description of the methodology pipeline used for the CNN architecture design. The suggested methodology is the Tree-structured Parzen Estimator (TPE) algorithm, a Bayesian optimization technique for automated architecture selection. By implementing TPE, we manage to explore and exploit efficiently a complex and large hyperparameter configuration space and automatically select optimal hyperparameters (such as number of convolutional layers, kernel size, strides, learning rates etc.). This automated approach reduces time consumption, errors, and the need for specialized knowledge in biology and biochemistry that would be associated with manual design. In addition to developing a pipeline for designing, training and evaluating a CNN for IMS data classification, we also apply a model agnostic interpretation methodology based on SHapley Additive exPlanations (SHAP) and provide SHAP score maps that visualize the importance of features in the spatial dimension of the IMS datacube.

In this thesis, we present and analyse the automated selection of 1-D CNN architectures for IMS data classification based on the TPE algorithm. Furthermore, we investigate a novel alternative to convolution, SCS, and evaluate its strengths and weaknesses in IMS data classification. The experimental results show that the TPE-generated CNN architectures outperform all the other applied classifiers. Finally, our interpretation of the CNN models reveals that accuracy performance alone might not be a sufficient criterion to trust the model's output. ...
Macro X-ray fluorescence (MA-XRF) is a recently developed technology allowing to obtain elemental information from cultural heritage objects. This information can, for example, be used to identify pigments used in a painting. Yet, the extended period of time it takes to scan an object is a major issue within MA-XRf. For instance, it took about 60 days to scan the Ghent Altarpiece. The long scanning time is a consequence of the necessary dwell time per pixel to create a robustly interpretable spectrum: the higher the dwell time, the higher the signalto-noise ratio (SNR), hence, the easier to detect elements. This thesis explores a possible solution for this problem using a denoising algorithm that increases the signal-to-noise ratio post-acquisition by exploiting the similarity between neighbouring pixels and spectra. To this end, a customized method of wavelet filter bank denoising is proposed. Current thresholding methods used in wavelet filter bank denoising are not suitable for filtering MA-XRF data, therefore, a novel thresholding method is introduced. Here, the widely used universal thresholding method is used as a basis, for which the formula for calculating the standard deviation of the detail coefficients of a channel is altered. Several design parameters of wavelet filter bank denoising were evaluated using a synthetic dataset, for which the performance quality indicators root mean square error (RMSE), mean absolute error (MAE) and SNR were determined. The parameters for which we optimized were the mother wavelet, the number of decomposition levels, and the number of neighbouring channels used for determining the standard deviation σ for thresholding. Good performance was obtained with the haar, db2, and coif1 wavelets, all at 3 levels of decomposition. A suitable number of neighbouring channels depended on the decomposition level and was determined to be 3 (on each side of the channel). Herewith, the signal-to-noise ratio was improved for both the average pixel spectra and the sum spectrum. The filtered synthetic dataset simulated to have a dwell time of 0.5 seconds had a SNR approximately equal to the raw synthetic dataset simulated to have a dwell time of 0.75 seconds. Hence, the algorithm succeeded in lowering the necessary dwell time. A case study of a daguerreotype was used to test the proposed denoising algorithm. ...
Master thesis (2021) - S.T. Jansen, R. van de Plas, L.G. Migas
Imaging Mass Spectrometry (IMS) is a powerful technique capable of extracting unlabeled spatial and chemical information from a biological tissue sample. Ever-increasing technological advancements have resulted in rapid growth of IMS data set sizes, scaling quadratically with the increasingly refined spatial resolution of its ion images. Dimensionality reduction techniques aim to make large data sets practically approachable by reducing the number of dimensions in the data while retaining as much information as possible. Nonlinear Dimensionality Reduction (NLDR) methods attempt to uncover an underlying nonlinear manifold or structure in the data by constructing a Low-Dimensional (LD) feature space with variables that are nonlinear combinations of the original features (i.e. mass-to-charge ratio (m/z) bins in IMS measurements). The t-Distributed Stochastic Neighbor Embedding (t-SNE) has been a common approach in NLDR methods, using force-directed graphs. Uniform Manifold Approximation and Projection (UMAP) works in a similar manner, but leans on a more versatile mathematical foundation in topology, is generally faster, and the method tends to scale better than t-SNE.

In this thesis, we introduce UMAPLUS, an extension of UMAP that takes not only spectral, but also spatial information into account. Our experiments show that UMAPLUS is capable of extracting the structure of a synthetic IMS data set better than UMAP. Besides UMAP’s standard spectral and random initializations, we introduce a new spatial initialization that provides more intuitive insight into the LD embedding relative to the spatial image domain. Standard UMAP entails several parameters that influence consistency, reliability, and quality of its LD embedding. Thus far, the influence of these parameters on IMS data sets has been barely investigated, and previous studies have consistently applied the default settings of UMAP. In this thesis, a Data-Driven UMAP (DD-UMAP) is constructed, which includes optimization of UMAP parameters in a data-driven manner. Several distance metrics are investigated, with cosine similarity providing the most robust results. A naive 1-D optimization procedure is compared to a multivariate Bayesian optimization approach, capable of optimizing multiple parameters simultaneously. Using an evaluation function partly based on the cost function of UMAP and the addition of spatial information, DD-UMAP is able to estimate and utilize an optimized input parameter set for UMAP in an unsupervised manner and on unlabeled IMS data sets. This is demonstrated on generated synthetic data, and two real world IMS data sets; a full mouse pup and a murine kidney. ...
Master thesis (2021) - M.D. Izarin, R. van de Plas, W.J. Niessen, D.H.J. Poot, Ribeiro Sabidussi
Recently, many advancements have been made in accelerated MRI reconstruction with the use of neural networks. Such deep learning methods learn a suitable MRI prior distribution from large sets of training data. For MRI images acquired with an uncommon scanning sequence, large datasets required for training are not available. Additionally, deep learning methods do not generalize well to unseen data. Therefore, in this research, a framework is proposed for accelerated MRI reconstruction trained with simulated data. The framework uses a Recurrent Inference Machine (RIM). The RIM is a deep learning framework that learns an iterative inference process. The RIM framework has been chosen as it is designed to learn the inversion process itself rather than the image statistics. Therefore, RIMs have a low tendency to overfit, and a high capacity to generalize to unseen data. The framework is evaluated by reconstructing undersampled data of in-vivo brain MRI images and comparing them with zero-filled reconstructions, reconstructions of an identical framework trained with in-vivo data and ESPIRiT reconstructions. The comparison shows that the framework does partly learn the inference process; however, the reconstructions still contain artefacts and the reconstruction of the framework trained with in-vivo data and the ESPIRiT method are of higher quality. For simulated data to replace in-vivo data for the training of the RIM, the simulated data has to be more similar to the in-vivo data. ...

Robust Matrix Decomposition for Spectral Imaging

Master thesis (2021) - R.A.R. Moens, R. van de Plas, B.R. Brandl
Modern imaging modalities across many application domains increasingly acquire a large number of very high-dimensional measurements, commonly collecting hundreds to millions of variables per spatial resolution element. That high-dimensional nature can severely challenge traditional (often Euclidean distance based) approaches to noise and dimensionality reduction. Furthermore, statistical analysis of such data is often hampered by the curse of dimensionality, concomitant with large numbers of pixels and channels, by the growing abundance of low signal-to-noise ratio (SNR) measurements, and by the detrimental effects of noise accumulation. It is therefore necessary to find efficient means of reducing high-dimensional (and large data size) measurement sets to a lower-dimensional representation while incurring minimal information loss. Addressing this challenge is essential (a) to enable processing of the massively multivariate measurement sets acquired by several promising imaging technologies today (commonly hundreds of gigabytes to terabytes per experiment), (b) to avoid that computational analysis becomes a bottleneck for the development of new instrumental capabilities, with hardware setups theoretically capable of yielding terabyte to petabyte imaging but currently considered impractical, and (c) to enable archiving massively multivariate measurement sets, often required to be stored for several years and containing information suitable for additional science projects. This thesis focuses on this challenge specifically. First, it explores structured and regularized matrix decomposition methods based on the l1-norm, e.g. building on principal component pursuit (PCP), to address this challenge. Second, it delivers custom implementations of these methods. Third, it develops and implements an application-driven framework for automatic setting of hyperparameters. Finally, it applies and compares these methods using several spectral imaging case studies spanning both very small scales, in molecular imaging of organic tissue, as well as very large scales, in the spectroscopic imaging of Outer space. ...

Class-aware Linear Feature Extraction in Imaging Mass Spectrometry

Retrieving actionable information from large datasets is increasingly computationally expensive due to the current trend of ever-increasing dataset sizes. Reducing dataset sizes with dimensionality reduction techniques is often necessary for statistical analysis techniques, such as classification, to be computationally feasible. Most dimensionality reduction methods do not require any additional information to accomplish their task. However, datasets used for classification, for example, are accompanied by a set of class-labels as well. This extra information can improve dimensionality reduction techniques by explicitly preserving features that explain differences between classes. A field where high-dimensional and large datasets are standard is Imaging Mass Spectrometry (IMS), a technique that simultaneously records the abundance and spatial location of molecules throughout biological tissue samples. Classification has been applied to IMS datasets for a wide range of scenarios, including the diagnosis of disease, distinguishing between tumour types for personalized treatment, and identifying biomarkers. A recently introduced dimensionality reduction method called Soft Discriminant Map (SDM), designed to incorporate class information and prevent overfitting when used on high-dimensional datasets, is a promising candidate to reduce the size and dimensionality of IMS datasets. However, SDM currently requires manual setting of a free parameter β that influences class separation in the newly constructed feature-space. This thesis explores the use of SDM on IMS datasets in classification use cases and proposes a framework to set β in a data-driven way: Data-Driven Soft Discriminant Map (DD-SDM). Furthermore, the sensitivity of the classification performance to changes in β is examined. DD-SDM is compared to similar state-of-the-art dimensionality reduction methods in terms of classification performance. The performed experiments show that DD-SDM successfully finds a value for β where the classification performance is on par with, or in some scenarios better than, state-of-the-art dimensionality reduction methods while using fewer features. Setting β either too low or too high results in a suboptimal feature space and worsens classification performance. Golden section search, the search strategy used to find the optimal β in DD-SDM, succeeds in finding the optimal β in fewer iterations than more naive methods. With the use of an artificial dataset in combination with a novel evaluation metric, the Peak Conservation Score (PCS), the distinctive ability of DD-SDM to discard features that are common between classes and to actively select for discriminative features is demonstrated. The DD-SDM framework is furthermore applied to real-world IMS measurements of rat brain and mouse kidney tissue. ...
Nowadays, video surveillance and motion detection system are widely used in various environments. With the relatively low-price cameras and highly automated monitoring system, video and image analysis on road, highway and skies becomes realistic. The key process in the analysis is to separate the useful information such as moving foreground objects from the original video sequence where Robust Principal Component Analysis (RPCA) plays an important role in extracting the foreground objects. RPCA have been widely used in data analysis and dimension reduction with applications in image recovery, information clustering and computer vision. But one drawback of RPCA lies in the fact that it does not guarantee the nonnegativity of pixels. It is important to have nonnegative foreground object since negative pixels that are not in the range between 0 and 255 are meaningless and the foreground objects are thus not visible. State-of-the-art methods do not consider the nonnegativity of the foreground object in their algorithms. This thesis focuses on the problem of extracting foreground moving object from background scenes and guarantee the nonnegativity of foreground object. This thesis proposes a method that combines RPCA and Nonnegative Matrix Factorization (NMF). It ensures the pixels that constitute the foreground object is nonnegative by using the basic model of RPCA and nonnegative components that NMF provides. The efficacy of the proposed algorithms is tested on publicly available dataset. Experiment shows in detail how the proposed algorithms achieve in recovering the foreground object with high true positive rate. Together with RPCA algorithm, the performance of recovery is compared and their advantages and disadvantages are discussed. ...
Master thesis (2020) - M. Bakos, R. van de Plas
In the field of biomedical imaging, images often report a combination of biologically induced variation, usually the goal of the imaging process (e.g. outlining an anatomical region or disease pattern), and non-biological variation, such as instrument or acquisition method-induced noise patterns.
Since some medical decisions are made based on imaging, separating the biological signal from noise is of significant importance (e.g. accelerates decision-making, reducing the chance of misdiagnosis).
Some non-biological variations that span a wide range of imaging modalities include e.g. viewport stitching artifacts, slice-to-slice interference, aliasing, and Gibbs-phenomena.
From a signal processing perspective, many of these can be modeled as quasiperiodic patterns.
Thus, removal of quasiperiodic patterns while preserving the underlying medical information is the main focus of this thesis.

Although in modern instruments, many forms of non-biological variation can be attenuated to be invisible to the naked eye, machine learning algorithms which are often used for classification of disease and segmentation of biological samples may be susceptible to even minor variations and noise patterns.
Development of entirely data-driven, unsupervised denoising techniques can potentially increase the effectiveness and reliability of such algorithms.
Furthermore, under certain image transformations, such as different color spaces, the Fourier and wavelet transform, and factorizations, such as principal component analysis and non-negative matrix factorization, as well as combinations of these, non-biological patterns can get amplified and become so prominent that much of the underlying biological information is concealed.
Removing these quasiperiodic patterns using current state-of-the-art algorithms still requires manual parameter tuning and prior expert knowledge, which is an impractical and possibly unnecessary expectation towards healthcare professionals.
The goal of this M.Sc. thesis is to develop an automated, data-driven framework, that is able to reliably identify, quantify, and eliminate quasiperiodic patterns within the images while retaining as much biological information as possible.

In this framework, named Quasiperiodic Image Denoising (QID), two novel algorithms are implemented, both operating in the Fourier domain.
One algorithm is based on robust principal component analysis (QID-RPCA) and the other uses the normalized median of absolute differences (QID-MADN).
The methods used to achieve unsupervised, data-driven denoising are described in detail.
This includes the use of histogram equalization for radial binning, an automated, sparsity-based approach to choosing the optimal aggregation level, and noise component attenuation based on radial frequency patterns.
The methodology is demonstrated through three case studies.
First, a synthetic dataset is used to compare the performance of the novel algorithms to the current state-of-the-art solutions.
Second, the performance is evaluated on two real-world datasets processed using a number of methods, e.g. factorization and different color spaces. One of these datasets is based on a microscopy image of a transversal section of a mouse brain and the other one is based on a microscopy image of a coronal section of a rat kidney.
Finally, a real-world, raw dataset is denoised consisting of a set of high-resolution fluorescent microscopy images of a human kidney.
Results indicate that the novel algorithms have higher denoising performance than previous approaches in the literature with notable improvements achieved for low-frequency corruptions. ...
The ability to locate specific objects within images is an essential step in various computer vision based engineering applications. Image segmentation is the task of dividing an image into "segments" that are uniform as well as homogeneous with respect to some characteristics, for example grey tone or texture as in Haralick et al. This thesis seeks to perform image segmentation using a Deep Learning (DL) approach in the area of warehouse automation, specifically focusing on an order picking use case of Vanderlande Industries (VI). Generally in literature, DL algorithms for image segmentation are split into two main classes: algorithms for RGB images and algorithms for RGB-D images. RGB stands for the Red, Green, and Blue values of a pixel. RGB-D stands for the Red, Green, Blue, and Depth values of a pixel. The depth value in this case differs from the RGB values in that it does not give a value for a colour intensity, but rather it gives a value for physical distance between the camera and the object it is capturing. The challenge addressed by this thesis focuses on whether the introduction of depth data results in a substantially better performance than using RGB-only images, based on a data-set provided by VI. Also, this thesis looks into the maximum allowed deviation along the X-axis in the registration of the depth data to the RGB images. Two networks from literature were investigated and implemented in MATLAB for this purpose: the SegNet architecture proposed by Badrinarayanan et al. and the FuseNet architecture proposed by Hazirbas et al. Through experiments we have found that, for this use case, the introduction of complementary depth data leads to an improvement over the use of RGB-only images. We also find that, for this use case, the maximum allowed deviation along the X-axis in the registration of the depth data to the RGB images is approximately equal to 1.67 millimetres. The results in this thesis seem to indicate that investing in acquiring an additional depth band does have a positive effect on the accuracy of image segmentation for order picking in warehouse automation. ...