Machine Learning for Iterative Strain Optimization

Doctoral Thesis (2026)
Author(s)

P.H. van Lent (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

Thomas Abeel – Promotor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

M.J.T. Reinders – Promotor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Pattern Recognition and Bioinformatics
DOI related publication
https://doi.org/10.4233/uuid:2c66d916-b8bf-4b1c-b3ec-2c44e5512b25 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Defense Date
03-07-2026
Awarding Institution
Delft University of Technology
Research Group
Pattern Recognition and Bioinformatics
Downloads counter
121
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Fermentation has played an important role in the production of foods such as bread, beer, wine, and cheese for thousands of years. These processes rely on micro-organisms that convert sugars into products such as alcohol, lactic acid, and carbon dioxide under oxygen-depleted conditions. Thanks to genetic modification, the variety of substances that can be produced through fermentation has expanded greatly. Fermentative production – also known as bioprocesses – offers a sustainable alternative to fossil fuel-based chemical production. However, metabolic engineering, in which the genetics of micro-organisms are specifically modified to produce desired metabolites in high yields, is highly complex, time-consuming, and costly. One of the main challenges is the enormous design space created by the complexity of cellular metabolism.

Machine learning can help explore this design space more efficiently, for example by predicting the performance of strain designs or suggesting new genetic modifications. In this thesis, we focus on improving these two applications. First, we develop a simulation tool that mimics metabolic processes, allowing us to compare different machine learning models and experimental strategies in a fair and cost-efficient manner. We then use the insights gained from these simulations to optimize yeast strains that produce p-Coumaric acid.

One drawback of many machine learning models is their limited transparency: it is often difficult to understand how a prediction is generated. In this thesis, we investigate how such models can be combined with mechanistic, mathematically formulated models. This hybrid approach brings together the predictive accuracy of machine learning and the interpretability of mechanistic models. We demonstrate that this integration results in more understandable models without sacrificing predictive performance.

Together, the methods and software developed in this thesis provide new tools for applying machine learning more effectively in metabolic engineering, with the aim of accelerating the development of sustainable bioprocesses.