M.J.T. Reinders
Please Note
102 records found
1
Approach. We present Augur, which predicts PAM50 subtype from a whole-slide image and localizes the tissue behind each call. Because slides are gigapixel with only slide-level labels, Augur adopts the standard weakly-supervised decomposition into two stages: a tile encoder first learns histology features through multi-task pretraining, then an attention-based multiple-instance-learning (MIL) aggregator pools tiles into a slide-level prediction. Our aggregator, DualCLAM, builds on CLAM, an existing model that scores tiles with per-class attention to form a slide representation and adds an instance-clustering loss that refines tile features from slide-level labels alone. DualCLAM augments it with auxiliary mutational-signature regression heads that contribute a molecular training signal and per-signature spatial explanations.
Results. On the TCGA-BRCA cohort under 5-fold stratified cross-validation, the encoder's pretext design dominates accuracy: it adds 0.077 macro AUROC over a segmentation-only baseline, roughly 5 times the effect of any MIL-side change. The best configuration, the four-pretext encoder paired with a CNV signature-regression head, reaches macro AUROC 0.7038. Never trained on annotations, the predicted-subtype attention maps align with pathologist-delineated tumor outlines on held-out slides. The auxiliary signature regression itself, however, is weak: the COSMIC exposure vectors are sparse, high-dimensional, and unevenly covered, ill-suited to the regression objective we used and likely better cast as multilabel classification. Even so, its attention heads still surface a small, stable subset of signatures (CN17, ID14, SBS33, SBS39).
Conclusions. H&E whole-slide images can yield interpretable, spatially-grounded subtype predictions at a level useful for triage. Accuracy hinges on the tile encoder's representation, so encoder pretraining warrants the most effort, but the aggregator should be redesigned, not abandoned: its weak signature branch points to two concrete fixes, predicting only a curated subset of signatures to shrink the decoder and recasting its objective from regression to multilabel classification. By pairing each prediction with visual and molecular evidence, Augur offers a low-cost, auditable pre-screen that prioritizes confirmatory molecular testing. ...
Approach. We present Augur, which predicts PAM50 subtype from a whole-slide image and localizes the tissue behind each call. Because slides are gigapixel with only slide-level labels, Augur adopts the standard weakly-supervised decomposition into two stages: a tile encoder first learns histology features through multi-task pretraining, then an attention-based multiple-instance-learning (MIL) aggregator pools tiles into a slide-level prediction. Our aggregator, DualCLAM, builds on CLAM, an existing model that scores tiles with per-class attention to form a slide representation and adds an instance-clustering loss that refines tile features from slide-level labels alone. DualCLAM augments it with auxiliary mutational-signature regression heads that contribute a molecular training signal and per-signature spatial explanations.
Results. On the TCGA-BRCA cohort under 5-fold stratified cross-validation, the encoder's pretext design dominates accuracy: it adds 0.077 macro AUROC over a segmentation-only baseline, roughly 5 times the effect of any MIL-side change. The best configuration, the four-pretext encoder paired with a CNV signature-regression head, reaches macro AUROC 0.7038. Never trained on annotations, the predicted-subtype attention maps align with pathologist-delineated tumor outlines on held-out slides. The auxiliary signature regression itself, however, is weak: the COSMIC exposure vectors are sparse, high-dimensional, and unevenly covered, ill-suited to the regression objective we used and likely better cast as multilabel classification. Even so, its attention heads still surface a small, stable subset of signatures (CN17, ID14, SBS33, SBS39).
Conclusions. H&E whole-slide images can yield interpretable, spatially-grounded subtype predictions at a level useful for triage. Accuracy hinges on the tile encoder's representation, so encoder pretraining warrants the most effort, but the aggregator should be redesigned, not abandoned: its weak signature branch points to two concrete fixes, predicting only a curated subset of signatures to shrink the decoder and recasting its objective from regression to multilabel classification. By pairing each prediction with visual and molecular evidence, Augur offers a low-cost, auditable pre-screen that prioritizes confirmatory molecular testing.
In this study, we investigate the relationships between these fragmentomic features in a genome-wide setting and evaluate their complementarity through multi-view intermediate integration for a binary classification task. Variance decomposition with correlation analysis showed that FSLR, MDS, and CNA capture partially non-redundant aspects of tumour-derived cfDNA signals, with only 7.2% overlap among outliers.
This biological complementarity did not translate into substantially improved predictive performance. The strongest downstream model was the PCA-based concatenation baseline using all three views, achieving an AUC of 0.969, with only marginal gains over CNA alone. In contrast, MOFA+ did not improve classification performance, reflecting a mismatch between its variance-maximization objective and cancer–healthy discrimination in cfDNA data. Similarly, Contrastive Multi-View Kernel Learning (CMK) failed to yield separable representations under an unsupervised objective, with meaningful class structure emerging only when supervision was introduced, yet still not surpassing the concatenation baseline.
Across all methods, CNA was consistently the most discriminative single feature, while FSLR provided an additional independent signal. MDS performed worst in all settings and contributed little to predictive performance. This limited contribution may reflect the chosen 5 Mb resolution rather than an inherent lack of biological signal, suggesting that feature-specific resolution optimization should be considered prior to integration. ...
In this study, we investigate the relationships between these fragmentomic features in a genome-wide setting and evaluate their complementarity through multi-view intermediate integration for a binary classification task. Variance decomposition with correlation analysis showed that FSLR, MDS, and CNA capture partially non-redundant aspects of tumour-derived cfDNA signals, with only 7.2% overlap among outliers.
This biological complementarity did not translate into substantially improved predictive performance. The strongest downstream model was the PCA-based concatenation baseline using all three views, achieving an AUC of 0.969, with only marginal gains over CNA alone. In contrast, MOFA+ did not improve classification performance, reflecting a mismatch between its variance-maximization objective and cancer–healthy discrimination in cfDNA data. Similarly, Contrastive Multi-View Kernel Learning (CMK) failed to yield separable representations under an unsupervised objective, with meaningful class structure emerging only when supervision was introduced, yet still not surpassing the concatenation baseline.
Across all methods, CNA was consistently the most discriminative single feature, while FSLR provided an additional independent signal. MDS performed worst in all settings and contributed little to predictive performance. This limited contribution may reflect the chosen 5 Mb resolution rather than an inherent lack of biological signal, suggesting that feature-specific resolution optimization should be considered prior to integration.
Predicting treatment of rheumatoid arthritis with LIVI
Prediction with a classifier on top of the LIVI model
Representation Learning for High-Dimensional Single-Cell Genomics with Variational Autoencoders
Using Associations Between Latent Factors and SNPs to Discover new eQTLs
changes in gene expression in that cell. This allows us to study the effect of genetics on diseases per cell instead of aggregated, since effects can differ per cell type. Traditional SNP to gene expression linking on the single-cell level suffers from the multiple testing burden, due to the great amount of SNPs and genes. To address this, a deep learning framework was developed recently to compress gene expression into low-dimensional encodings and reconstruct the gene expression linearly from these encodings, enabling direct interpretation of the latent space. This model is called Latent Interaction Variational Inference (LIVI). Here, we determine whether the latent factors of this model can serve as a quantitative trait for Single Nucleotide Polymorphisms (SNPs) that associate with Rheumatoid Arthritis (RA) on a dataset with RA patients. RA is a chronic disease characterized by progressive damage of the joints. In this study, we found 617 out of 700 latent factors correlating to at least one SNP, using a linear mixed model. We also found that genes that are associated with RA in a Genome Wide Association Study have a higher loading for associated SNP-Latent factor pairs then for none associated one. We also identified genes affected by GWAS-identified risk SNPs for which the original GWAS did not identify a functionally associated gene. We conclude that the latent factors of the LIVI model can be used as a quantitative trait for SNPs, and used these latent factors to discover trans-eQTLs. ...
changes in gene expression in that cell. This allows us to study the effect of genetics on diseases per cell instead of aggregated, since effects can differ per cell type. Traditional SNP to gene expression linking on the single-cell level suffers from the multiple testing burden, due to the great amount of SNPs and genes. To address this, a deep learning framework was developed recently to compress gene expression into low-dimensional encodings and reconstruct the gene expression linearly from these encodings, enabling direct interpretation of the latent space. This model is called Latent Interaction Variational Inference (LIVI). Here, we determine whether the latent factors of this model can serve as a quantitative trait for Single Nucleotide Polymorphisms (SNPs) that associate with Rheumatoid Arthritis (RA) on a dataset with RA patients. RA is a chronic disease characterized by progressive damage of the joints. In this study, we found 617 out of 700 latent factors correlating to at least one SNP, using a linear mixed model. We also found that genes that are associated with RA in a Genome Wide Association Study have a higher loading for associated SNP-Latent factor pairs then for none associated one. We also identified genes affected by GWAS-identified risk SNPs for which the original GWAS did not identify a functionally associated gene. We conclude that the latent factors of the LIVI model can be used as a quantitative trait for SNPs, and used these latent factors to discover trans-eQTLs.
Capturing Clinical Heterogeneity in Rheumatoid Arthritis
Evaluating the LIVI Latent Space using Gene Expression Data
This paper investigates whether integrating chromatin accessibility into GRN inference can help explain gene regulatory changes in specific brain cell types during AD progression by evaluating whether ScReNI's ATAC-derived terms improve biological support beyond RNA-only inference. Using a Python reimplementation, we decompose the published ScReNI formula and assess formula variants on the mouse retina development benchmark with ChIP-Atlas precision and network-clustering ARI. We then apply the same component analysis to SEA-AD MTG data and evaluate whether literature reported AD-related microglia genes survive feature selection.
The results suggest that removing the regulator-locus term and using TF-specific target-peak attribution improves performance of the ScReNI weight formula on the mouse retina development dataset. However, before ScReNI can support strong AD-specific regulatory claims in SEA-AD, feature selection must retain disease-relevant regulators and validation must be performed using human brain cell-type-specific benchmarks. ...
This paper investigates whether integrating chromatin accessibility into GRN inference can help explain gene regulatory changes in specific brain cell types during AD progression by evaluating whether ScReNI's ATAC-derived terms improve biological support beyond RNA-only inference. Using a Python reimplementation, we decompose the published ScReNI formula and assess formula variants on the mouse retina development benchmark with ChIP-Atlas precision and network-clustering ARI. We then apply the same component analysis to SEA-AD MTG data and evaluate whether literature reported AD-related microglia genes survive feature selection.
The results suggest that removing the regulator-locus term and using TF-specific target-peak attribution improves performance of the ScReNI weight formula on the mouse retina development dataset. However, before ScReNI can support strong AD-specific regulatory claims in SEA-AD, feature selection must retain disease-relevant regulators and validation must be performed using human brain cell-type-specific benchmarks.
Multi-omic Latent Interaction Modelling at Single-Cell Resolution
Extending Latent Interaction Variational Inference (LIVI) Model with Protein Modality
Associating Single-Cell Latent Factors with Genetic Risk
An analysis on a clinical patient cohort
Analysis of HVG use in the ScReNI pipeline
A comparison of global and cell-type specific HVG selection
Specifically, this work compares the highly variable gene (HVG) selection strategy used in ScReNI with a newly proposed approach: type-specific HVG selection. Instead of selecting the top HVGs globally from the entire dataset, we propose selecting the top HVGs within each cell type and inferring the GRN of each cell using only genes specific to its cell type. The comparison is conducted across multiple structural and biological metrics, and the type-specific selection approach shows overall improved performance compared to the original method. ...
Specifically, this work compares the highly variable gene (HVG) selection strategy used in ScReNI with a newly proposed approach: type-specific HVG selection. Instead of selecting the top HVGs globally from the entire dataset, we propose selecting the top HVGs within each cell type and inferring the GRN of each cell using only genes specific to its cell type. The comparison is conducted across multiple structural and biological metrics, and the type-specific selection approach shows overall improved performance compared to the original method.
Cell-specific Gene Regulatory Networks in Alzheimer’s Disease
A Differential Analysis of Regulatory Changes Across Cell Types
cell’s dysfunction is thought to involve changes in how its genes are regulated rather
than only changes in their expression. Gene regulatory networks (GRNs) model these
regulatory decisions, and recent single-cell methods can infer a separate network for
each individual cell. How such cell-specific networks change in disease, and whether
any changes are shared across cell types or unique to particular ones, remains largely
unexplored. This work asks which regulatory relationships differ between AD and
healthy cells, and to what extent these differences are cell-type specific. To address
this, the cell-specific GRN method ScReNI is reimplemented in Python and applied
to paired single-nucleus RNA and ATAC data from the SEA-AD cohort (27 donors),
inferring one network per cell for microglia and an excitatory-neuron subclass (L2/3
IT). A differential pipeline then compares the networks against disease severity at
the level of individual transcription-factor-to-target edges and of co-regulated gene
modules, complemented by a module-preservation test, treating the donor as the unit of replication. The inferred networks change relatively little with disease: no edge survives stringent correction and weight-based filtering, the module co-regulation structure is preserved, and no module-specific shift is detected in either cell type. The one robust signal is a modest, severity-graded decline in overall regulatory activity in L2/3 IT neurons that persists after adjusting for sequencing depth and is absent in microglia. The results are best read as an absence of strong evidence rather than evidence of absence, motivating larger cohorts and broader gene panels in future work. ...
cell’s dysfunction is thought to involve changes in how its genes are regulated rather
than only changes in their expression. Gene regulatory networks (GRNs) model these
regulatory decisions, and recent single-cell methods can infer a separate network for
each individual cell. How such cell-specific networks change in disease, and whether
any changes are shared across cell types or unique to particular ones, remains largely
unexplored. This work asks which regulatory relationships differ between AD and
healthy cells, and to what extent these differences are cell-type specific. To address
this, the cell-specific GRN method ScReNI is reimplemented in Python and applied
to paired single-nucleus RNA and ATAC data from the SEA-AD cohort (27 donors),
inferring one network per cell for microglia and an excitatory-neuron subclass (L2/3
IT). A differential pipeline then compares the networks against disease severity at
the level of individual transcription-factor-to-target edges and of co-regulated gene
modules, complemented by a module-preservation test, treating the donor as the unit of replication. The inferred networks change relatively little with disease: no edge survives stringent correction and weight-based filtering, the module co-regulation structure is preserved, and no module-specific shift is detected in either cell type. The one robust signal is a modest, severity-graded decline in overall regulatory activity in L2/3 IT neurons that persists after adjusting for sequencing depth and is absent in microglia. The results are best read as an absence of strong evidence rather than evidence of absence, motivating larger cohorts and broader gene panels in future work.
The Geography of Gene Regulation in Alzheimer's Disease
Inferring and Analysing Spatial Gene Regulatory Networks with ScReNI and Tangram
To address this problem, the influence of varying amounts of additional data on generalisation to novel cloud regimes is evaluated. The study compares representation learning strategies, including self-supervised and supervised approaches, to assess their effectiveness in structuring the latent space for distinguishing cloud types. In parallel, two learning paradigms, transfer learning and episodic meta-learning, are analysed to determine how effectively they incorporate additional data when adapting to novel classes.
The results show that self-supervised learning is most effective in the strict few-shot regime, while supervised transfer learning makes the most effective use of additional data. In particular, Barlow Twins achieves the strongest performance under minimal data by avoiding reliance on noisy labels. When additional out-of-domain data, such as ImageNet and auxiliary cloud datasets, is introduced, supervised pre-training combined with transfer learning attains performance comparable to standard supervised learning, while requiring only a small number of labelled examples. ...
To address this problem, the influence of varying amounts of additional data on generalisation to novel cloud regimes is evaluated. The study compares representation learning strategies, including self-supervised and supervised approaches, to assess their effectiveness in structuring the latent space for distinguishing cloud types. In parallel, two learning paradigms, transfer learning and episodic meta-learning, are analysed to determine how effectively they incorporate additional data when adapting to novel classes.
The results show that self-supervised learning is most effective in the strict few-shot regime, while supervised transfer learning makes the most effective use of additional data. In particular, Barlow Twins achieves the strongest performance under minimal data by avoiding reliance on noisy labels. When additional out-of-domain data, such as ImageNet and auxiliary cloud datasets, is introduced, supervised pre-training combined with transfer learning attains performance comparable to standard supervised learning, while requiring only a small number of labelled examples.
Improving Remote Cardiovascular Care with Wearable Data
Algorithms, Study Design, and Subject-Specific Adaptation
By assessing how far wearable-based research has progressed toward operational deployment and identify critical shortcomings in real-world utility and generalizability, we confront several major challenges intrinsic to this domain: the medical interpretability of noisy consumer-grade signals, high inter-subject variability, and the inherent complexity of timeseries data that varies with context (e.g., day/night cycles, physical activity).
Our solution strategy is grounded in machine learning techniques that aim to learn robust, transferable representations of physiological data. In particular, we explore contrastive learning, weak supervision, and morphological modeling—such as acceleration-deceleration curve analysis— as tools to extract clinically relevant patterns. These methods are evaluated across both publicly available and proprietary datasets to ensure applicability to diverse populations.
By addressing these challenges, this dissertation advances the case for smartwatches as viable tools for longitudinal, data-efficient cardiovascular monitoring, contributing to a future in which early detection of conditions like atrial fibrillation and heart failure is feasible at scale in everyday settings.
...
By assessing how far wearable-based research has progressed toward operational deployment and identify critical shortcomings in real-world utility and generalizability, we confront several major challenges intrinsic to this domain: the medical interpretability of noisy consumer-grade signals, high inter-subject variability, and the inherent complexity of timeseries data that varies with context (e.g., day/night cycles, physical activity).
Our solution strategy is grounded in machine learning techniques that aim to learn robust, transferable representations of physiological data. In particular, we explore contrastive learning, weak supervision, and morphological modeling—such as acceleration-deceleration curve analysis— as tools to extract clinically relevant patterns. These methods are evaluated across both publicly available and proprietary datasets to ensure applicability to diverse populations.
By addressing these challenges, this dissertation advances the case for smartwatches as viable tools for longitudinal, data-efficient cardiovascular monitoring, contributing to a future in which early detection of conditions like atrial fibrillation and heart failure is feasible at scale in everyday settings.
Safer causal inference
Theory and algorithms for falsification, trial augmentation and policy evaluation
In Part One, we address the first aspect of detecting violations of causal identification assumptions. We focus on settings with data from multiple sources, such as hospitals or locations, where distributional shifts naturally occur. Under specific independence conditions on the causal mechanisms driving these shifts, we first present a nonparametric test to falsify the assumption of no unmeasured confounding. To obtain these results, we introduce a novel technique utilizing hierarchical causal graphical models. Thereafter, we focus on improving the statistical efficiency of this test, which is achieved by reformulating the independence condition using parameterized linear models. Finally, we extend the hierarchical modeling approach to other identification settings, specifically by testing the validity of mediators and instrumental variables used in two additional common identification strategies.
In Parts Two and Three, we develop methods that instead are robust when causal identification assumptions are violated.We revisit two commonly occurring problem settings when doing causal inference and demonstrate that it is possible to develop methods that either remove the need for, or rely on, weaker and more plausible assumptions than those traditionally made. In the first setting, we study the problemof augmenting randomized trials using external data to improve efficiency in treatment effect estimation. Typically, such approaches rely on a transportability assumption that relate the populations underlying the trial and external data. But when this transportability assumption is violated, integrating external data can introduce substantial bias. To address this, we propose a novel and efficient estimator that incorporates external data and show that this estimator improves inference on the average treatment effect while guaranteeing that it never performs worse, and sometimes performs better, than the estimator that relies solely on trial data.We further adapt this estimator to learn heterogeneous treatment effects within the trial population and show that similar safety guarantees hold for this problem.
In the second setting, we examine the evaluation of treatment allocation strategies using Qini curves. Standard methods for estimating Qini curves assume no interference between treated units, meaning that the treatment of one unit does not affect others. However, when interference is present, these Qini curves can be misleading and lead to incorrect evaluation of treatment allocation strategies.We therefore propose multiple estimators to handle the interference, specifically in settings where units within a cluster may affect one another but not units in other clusters.We identify a bias-variance trade-off in these estimators and, through both theoretical and empirical results, provide practical guidance on how practitioners can choose among them. The dissertation concludes with a discussion of broader considerations, limitations of the presented research, and potential directions for future work.We find that it is indeed possible to make causal inference safer by detecting assumption violations and reducing reliance on untestable assumptions. Nonetheless, many open and important questions remain, offering promising avenues for further research on this topic. ...
In Part One, we address the first aspect of detecting violations of causal identification assumptions. We focus on settings with data from multiple sources, such as hospitals or locations, where distributional shifts naturally occur. Under specific independence conditions on the causal mechanisms driving these shifts, we first present a nonparametric test to falsify the assumption of no unmeasured confounding. To obtain these results, we introduce a novel technique utilizing hierarchical causal graphical models. Thereafter, we focus on improving the statistical efficiency of this test, which is achieved by reformulating the independence condition using parameterized linear models. Finally, we extend the hierarchical modeling approach to other identification settings, specifically by testing the validity of mediators and instrumental variables used in two additional common identification strategies.
In Parts Two and Three, we develop methods that instead are robust when causal identification assumptions are violated.We revisit two commonly occurring problem settings when doing causal inference and demonstrate that it is possible to develop methods that either remove the need for, or rely on, weaker and more plausible assumptions than those traditionally made. In the first setting, we study the problemof augmenting randomized trials using external data to improve efficiency in treatment effect estimation. Typically, such approaches rely on a transportability assumption that relate the populations underlying the trial and external data. But when this transportability assumption is violated, integrating external data can introduce substantial bias. To address this, we propose a novel and efficient estimator that incorporates external data and show that this estimator improves inference on the average treatment effect while guaranteeing that it never performs worse, and sometimes performs better, than the estimator that relies solely on trial data.We further adapt this estimator to learn heterogeneous treatment effects within the trial population and show that similar safety guarantees hold for this problem.
In the second setting, we examine the evaluation of treatment allocation strategies using Qini curves. Standard methods for estimating Qini curves assume no interference between treated units, meaning that the treatment of one unit does not affect others. However, when interference is present, these Qini curves can be misleading and lead to incorrect evaluation of treatment allocation strategies.We therefore propose multiple estimators to handle the interference, specifically in settings where units within a cluster may affect one another but not units in other clusters.We identify a bias-variance trade-off in these estimators and, through both theoretical and empirical results, provide practical guidance on how practitioners can choose among them. The dissertation concludes with a discussion of broader considerations, limitations of the presented research, and potential directions for future work.We find that it is indeed possible to make causal inference safer by detecting assumption violations and reducing reliance on untestable assumptions. Nonetheless, many open and important questions remain, offering promising avenues for further research on this topic.
Computer vision concerns itself with the research and development of deep learning models that work on visual data. These vision models are already heavily integrated into society, powering real-world applications such as automated radiology in hospitals, self-driving cars, and autonomous drones. However, it takes a lot of data, in the form of datasets containing thousands or millions of images, to learn reliable vision models. This thesis explores the role that spatial biases (prior knowledge on the position and pose of objects in the image) can play in learning better and more data-efficient vision models.
We find that the practice of integrating prior knowledge on spatial biases (inductive spatial biases) can help to learn biases that are otherwise hard or impossible to learn. Though inductive bias can be difficult and time-consuming to design, and often increases inference cost, integrating inductive bias can result in better performance and greater data efficiency. This work showcases these patterns in spatial biases, specifically position bias and scale bias.
We find that position bias may be learned to some degree by models without the proper inductive bias, but that inductive bias helps to model these biases and improves performance. We show that whether learning position bias is helpful depends on the data. We contribute measures for position bias in vision models in general, as well as in Vision Transformers specifically, to enable the discovery of these findings. We propose an inductive bias on the position embedding of ViTs to better (un)learn position bias.
For scale bias, we find that existing scale-equivariant models for scale bias need to be tuned to the scale distribution of the data. We propose an inductive bias that allows scale-equivariant models to learn the scale bias of the dataset, thereby fitting the data better. We also propose an alternative parameterization of convolutions called MAGNet that can be adapted to known scale distributions present in the data. Models using MAGNets (FlexNets) can be much shallower and do not require pooling.
There are those who advocate against spending much time on inductive biases. The “bitter lesson” of Richard Sutton prescribes that we should simply add more data, not more inductive bias. However, data will run out at some point, perhaps sooner rather than later. Besides raw performance of vision models, given as much data as possible, should not be our only goal: data-deficient settings are real, plentiful, and important. Data-efficient vision models are the future of our field, and the search for appropriate inductive biases will remain an important endeavor. ...
Computer vision concerns itself with the research and development of deep learning models that work on visual data. These vision models are already heavily integrated into society, powering real-world applications such as automated radiology in hospitals, self-driving cars, and autonomous drones. However, it takes a lot of data, in the form of datasets containing thousands or millions of images, to learn reliable vision models. This thesis explores the role that spatial biases (prior knowledge on the position and pose of objects in the image) can play in learning better and more data-efficient vision models.
We find that the practice of integrating prior knowledge on spatial biases (inductive spatial biases) can help to learn biases that are otherwise hard or impossible to learn. Though inductive bias can be difficult and time-consuming to design, and often increases inference cost, integrating inductive bias can result in better performance and greater data efficiency. This work showcases these patterns in spatial biases, specifically position bias and scale bias.
We find that position bias may be learned to some degree by models without the proper inductive bias, but that inductive bias helps to model these biases and improves performance. We show that whether learning position bias is helpful depends on the data. We contribute measures for position bias in vision models in general, as well as in Vision Transformers specifically, to enable the discovery of these findings. We propose an inductive bias on the position embedding of ViTs to better (un)learn position bias.
For scale bias, we find that existing scale-equivariant models for scale bias need to be tuned to the scale distribution of the data. We propose an inductive bias that allows scale-equivariant models to learn the scale bias of the dataset, thereby fitting the data better. We also propose an alternative parameterization of convolutions called MAGNet that can be adapted to known scale distributions present in the data. Models using MAGNets (FlexNets) can be much shallower and do not require pooling.
There are those who advocate against spending much time on inductive biases. The “bitter lesson” of Richard Sutton prescribes that we should simply add more data, not more inductive bias. However, data will run out at some point, perhaps sooner rather than later. Besides raw performance of vision models, given as much data as possible, should not be our only goal: data-deficient settings are real, plentiful, and important. Data-efficient vision models are the future of our field, and the search for appropriate inductive biases will remain an important endeavor.
Machine learning can help explore this design space more efficiently, for example by predicting the performance of strain designs or suggesting new genetic modifications. In this thesis, we focus on improving these two applications. First, we develop a simulation tool that mimics metabolic processes, allowing us to compare different machine learning models and experimental strategies in a fair and cost-efficient manner. We then use the insights gained from these simulations to optimize yeast strains that produce p-Coumaric acid.
One drawback of many machine learning models is their limited transparency: it is often difficult to understand how a prediction is generated. In this thesis, we investigate how such models can be combined with mechanistic, mathematically formulated models. This hybrid approach brings together the predictive accuracy of machine learning and the interpretability of mechanistic models. We demonstrate that this integration results in more understandable models without sacrificing predictive performance.
Together, the methods and software developed in this thesis provide new tools for applying machine learning more effectively in metabolic engineering, with the aim of accelerating the development of sustainable bioprocesses. ...
Machine learning can help explore this design space more efficiently, for example by predicting the performance of strain designs or suggesting new genetic modifications. In this thesis, we focus on improving these two applications. First, we develop a simulation tool that mimics metabolic processes, allowing us to compare different machine learning models and experimental strategies in a fair and cost-efficient manner. We then use the insights gained from these simulations to optimize yeast strains that produce p-Coumaric acid.
One drawback of many machine learning models is their limited transparency: it is often difficult to understand how a prediction is generated. In this thesis, we investigate how such models can be combined with mechanistic, mathematically formulated models. This hybrid approach brings together the predictive accuracy of machine learning and the interpretability of mechanistic models. We demonstrate that this integration results in more understandable models without sacrificing predictive performance.
Together, the methods and software developed in this thesis provide new tools for applying machine learning more effectively in metabolic engineering, with the aim of accelerating the development of sustainable bioprocesses.
In this context, the thesis takes a step back and asks a more fundamental question: how can we reliably reason about generalization when data is scarce and the behavior of learning curves is itself uncertain? Rather than treating learning curves as simple, and monotonic functions, we study their full statistical structure. We show that variability across training subsets can influence model comparison, decision making, and performance extrapolation. In addition, we investigate conditions under which monotonic improvement can be guaranteed or encouraged. Beyond single task learning, we also examine meta-learning, where information from multiple related tasks is leveraged to improve generalization performance while reducing the amount of data required from any individual task.
We begin by showing that the mean, as a statistical summary of learning curves, may not provide a reliable estimate of performance. We demonstrate that generalization performance distributions are often skewed and heavy tailed, regardless of how they are obtained. As a result, relying solely on the mean for model selection can be suboptimal for some problems.
Next, we propose a semi parametric extrapolation method that adapts its inductive bias to capture complex and potentially non monotonic patterns. This approach improves predictive reliability in settings where additional data collection is costly or infeasible and where learning curves may not exhibit monotonic behavior.
We then study the monotonicity of learning curves under specific conditions. For linear regression, we show that a single gradient update is sufficient to ensure monotonic improvement, provided that the learning rate does not exceed a certain threshold. To construct similarly monotonic learners in practice, we propose a data driven approach for selecting both the learning rate and the initial parameter estimates.
Finally, we investigate the learning curves of a meta learning algorithm. Through controlled synthetic experiments, we analyze the generalization performance of both meta learners and task specific learners, providing insights into how properties of the task distribution influence generalization under a limited adaptation stage consisting of a single gradient update.
...
In this context, the thesis takes a step back and asks a more fundamental question: how can we reliably reason about generalization when data is scarce and the behavior of learning curves is itself uncertain? Rather than treating learning curves as simple, and monotonic functions, we study their full statistical structure. We show that variability across training subsets can influence model comparison, decision making, and performance extrapolation. In addition, we investigate conditions under which monotonic improvement can be guaranteed or encouraged. Beyond single task learning, we also examine meta-learning, where information from multiple related tasks is leveraged to improve generalization performance while reducing the amount of data required from any individual task.
We begin by showing that the mean, as a statistical summary of learning curves, may not provide a reliable estimate of performance. We demonstrate that generalization performance distributions are often skewed and heavy tailed, regardless of how they are obtained. As a result, relying solely on the mean for model selection can be suboptimal for some problems.
Next, we propose a semi parametric extrapolation method that adapts its inductive bias to capture complex and potentially non monotonic patterns. This approach improves predictive reliability in settings where additional data collection is costly or infeasible and where learning curves may not exhibit monotonic behavior.
We then study the monotonicity of learning curves under specific conditions. For linear regression, we show that a single gradient update is sufficient to ensure monotonic improvement, provided that the learning rate does not exceed a certain threshold. To construct similarly monotonic learners in practice, we propose a data driven approach for selecting both the learning rate and the initial parameter estimates.
Finally, we investigate the learning curves of a meta learning algorithm. Through controlled synthetic experiments, we analyze the generalization performance of both meta learners and task specific learners, providing insights into how properties of the task distribution influence generalization under a limited adaptation stage consisting of a single gradient update.