Multi-View Feature Analysis of Cell-Free DNA Data for Cancer Prediction

Master Thesis (2026)
Author(s)

L. Navarčíková (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

M.J.T. Reinders – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

I.B. Pronk – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

W.P. Brinkman – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
29-06-2026
Awarding Institution
Delft University of Technology
Programme
Computer Science, Data Science and Artificial Intelligence Technology
Faculty
Electrical Engineering, Mathematics and Computer Science
Downloads counter
23
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Cell-free DNA (cfDNA) fragmentomics using liquid biopsy data has emerged as a promising minimally invasive approach for cancer detection. While multiple fragmentomic features, including the Fragment Short Long Ratio (FSLR), Motif Diversity Score (MDS), and Copy Number Alterations (CNA), have individually demonstrated predictive value, their complementarity and optimal integration remain insufficiently explored.

In this study, we investigate the relationships between these fragmentomic features in a genome-wide setting and evaluate their complementarity through multi-view intermediate integration for a binary classification task. Variance decomposition with correlation analysis showed that FSLR, MDS, and CNA capture partially non-redundant aspects of tumour-derived cfDNA signals, with only 7.2% overlap among outliers.

This biological complementarity did not translate into substantially improved predictive performance. The strongest downstream model was the PCA-based concatenation baseline using all three views, achieving an AUC of 0.969, with only marginal gains over CNA alone. In contrast, MOFA+ did not improve classification performance, reflecting a mismatch between its variance-maximization objective and cancer–healthy discrimination in cfDNA data. Similarly, Contrastive Multi-View Kernel Learning (CMK) failed to yield separable representations under an unsupervised objective, with meaningful class structure emerging only when supervision was introduced, yet still not surpassing the concatenation baseline.

Across all methods, CNA was consistently the most discriminative single feature, while FSLR provided an additional independent signal. MDS performed worst in all settings and contributed little to predictive performance. This limited contribution may reflect the chosen 5 Mb resolution rather than an inherent lack of biological signal, suggesting that feature-specific resolution optimization should be considered prior to integration.

Files

Thesis_Final.pdf
(pdf | 18.3 Mb)
License info not available