GB
G.N.J.C. Bierkens
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
34 records found
1
In short Theory of Mind (ToM) is the ability to think about what someone else is thinking, with the added complexity of knowing that others might think differently than ourselves. On a more practical level this translates to choosing actions while keeping in mind what others are thinking about. These actions have to serve two purposes, the first one is to take advantage from this ability (exploitation), the second is to help us gather information to refine our knowledge about others (exploration). In the case where the dynamics unfold in a Partially Observable Markov Decision Process (POMDP), a framework that evaluates the cost of actions both from their exploratory, as well as their exploitative benefits, is the Active Inference Framework (AIF), whose formulas rely on an observation model P(o|s) and a preference distribution over states P(s).
At the moment, the area of the AIF dealing with cost of actions that keep others in mind is not fully developed and furthermore methods to learn the observation model and the preferences of others, so as to be able to model others over time, are still missing.
In this thesis we restricted ourselves to the case of a POMDP with finite number of states, observations and actions and only modeled the situation for two agents. We started by generalizing the overall formulation of the AIF by using beliefs about the observation model and the preferences of the other so as to encode the uncertainty we have about those quantities. We then extended the current theory of the AIF by proposing an action cost function which evaluates the cost of our actions while considering the presumed (re)actions of the other. Furthermore we developed methods to update the belief about the observation model and the preferences of the other, by using the actions the we see it take and the observations we see it receive. Finally we tested the central formulas of the AIF on a toy example to see if they work and also tested our function in a toy experiment, both in the case where the observation model of the other is given as well as in the case where it has to be learned. ...
At the moment, the area of the AIF dealing with cost of actions that keep others in mind is not fully developed and furthermore methods to learn the observation model and the preferences of others, so as to be able to model others over time, are still missing.
In this thesis we restricted ourselves to the case of a POMDP with finite number of states, observations and actions and only modeled the situation for two agents. We started by generalizing the overall formulation of the AIF by using beliefs about the observation model and the preferences of the other so as to encode the uncertainty we have about those quantities. We then extended the current theory of the AIF by proposing an action cost function which evaluates the cost of our actions while considering the presumed (re)actions of the other. Furthermore we developed methods to update the belief about the observation model and the preferences of the other, by using the actions the we see it take and the observations we see it receive. Finally we tested the central formulas of the AIF on a toy example to see if they work and also tested our function in a toy experiment, both in the case where the observation model of the other is given as well as in the case where it has to be learned. ...
In short Theory of Mind (ToM) is the ability to think about what someone else is thinking, with the added complexity of knowing that others might think differently than ourselves. On a more practical level this translates to choosing actions while keeping in mind what others are thinking about. These actions have to serve two purposes, the first one is to take advantage from this ability (exploitation), the second is to help us gather information to refine our knowledge about others (exploration). In the case where the dynamics unfold in a Partially Observable Markov Decision Process (POMDP), a framework that evaluates the cost of actions both from their exploratory, as well as their exploitative benefits, is the Active Inference Framework (AIF), whose formulas rely on an observation model P(o|s) and a preference distribution over states P(s).
At the moment, the area of the AIF dealing with cost of actions that keep others in mind is not fully developed and furthermore methods to learn the observation model and the preferences of others, so as to be able to model others over time, are still missing.
In this thesis we restricted ourselves to the case of a POMDP with finite number of states, observations and actions and only modeled the situation for two agents. We started by generalizing the overall formulation of the AIF by using beliefs about the observation model and the preferences of the other so as to encode the uncertainty we have about those quantities. We then extended the current theory of the AIF by proposing an action cost function which evaluates the cost of our actions while considering the presumed (re)actions of the other. Furthermore we developed methods to update the belief about the observation model and the preferences of the other, by using the actions the we see it take and the observations we see it receive. Finally we tested the central formulas of the AIF on a toy example to see if they work and also tested our function in a toy experiment, both in the case where the observation model of the other is given as well as in the case where it has to be learned.
At the moment, the area of the AIF dealing with cost of actions that keep others in mind is not fully developed and furthermore methods to learn the observation model and the preferences of others, so as to be able to model others over time, are still missing.
In this thesis we restricted ourselves to the case of a POMDP with finite number of states, observations and actions and only modeled the situation for two agents. We started by generalizing the overall formulation of the AIF by using beliefs about the observation model and the preferences of the other so as to encode the uncertainty we have about those quantities. We then extended the current theory of the AIF by proposing an action cost function which evaluates the cost of our actions while considering the presumed (re)actions of the other. Furthermore we developed methods to update the belief about the observation model and the preferences of the other, by using the actions the we see it take and the observations we see it receive. Finally we tested the central formulas of the AIF on a toy example to see if they work and also tested our function in a toy experiment, both in the case where the observation model of the other is given as well as in the case where it has to be learned.
The Mpemba effect refers to the phenomenon that a system prepared at a higher initial temperature may reach thermal equilibrium faster than an otherwise identical system prepared at a lower temperature. In this thesis, the Mpemba effect is investigated for the Glauber dynamics of the Ising model. Using a spectral decomposition of the transition rate matrix, a coefficient a2(T) is derived whose behaviour determines whether the spectral criterion for the Mpemba effect is satisfied. In addition, the role of metastability is examined through a reduced description of the dynamics based on transitions between metastable wells. Numerical experiments are performed on finite one-dimensional Ising systems to investigate the behaviour of a2(T). Several examples exhibiting non-monotonic behaviour of a2(T) are found, indicating the possible occurrence of the Mpemba effect. However, many of these examples do not possess a clear metastable structure, and the reduced metastable dynamics often fail to reproduce the behaviour observed in the full system. These results suggest that the Mpemba effect is not solely determined by metastability, but depends more generally on the spectral structure of the dynamics and on how thermal initial states interact with the slow relaxation modes.
...
The Mpemba effect refers to the phenomenon that a system prepared at a higher initial temperature may reach thermal equilibrium faster than an otherwise identical system prepared at a lower temperature. In this thesis, the Mpemba effect is investigated for the Glauber dynamics of the Ising model. Using a spectral decomposition of the transition rate matrix, a coefficient a2(T) is derived whose behaviour determines whether the spectral criterion for the Mpemba effect is satisfied. In addition, the role of metastability is examined through a reduced description of the dynamics based on transitions between metastable wells. Numerical experiments are performed on finite one-dimensional Ising systems to investigate the behaviour of a2(T). Several examples exhibiting non-monotonic behaviour of a2(T) are found, indicating the possible occurrence of the Mpemba effect. However, many of these examples do not possess a clear metastable structure, and the reduced metastable dynamics often fail to reproduce the behaviour observed in the full system. These results suggest that the Mpemba effect is not solely determined by metastability, but depends more generally on the spectral structure of the dynamics and on how thermal initial states interact with the slow relaxation modes.
Prior specification is a fundamental challenge in Bayesian inference. Traditionally, a prior represents a belief over the behavior of a statistical model based on expert knowledge. Such knowledge is not always available, or can be hard to express in a numerical form. As an answer, this thesis introduces the Prior-Learning Prior-Fitted Network (PLPFN), a transformer-based meta-learning model that amortizes prior learning. The model takes in a set of related tasks and outputs fully Bayesian uncertainty estimates over prior parameters in a single forward pass. The novelty lies in providing interpretable prior parameter outputs and examining the prior learning problem through the lens of hierarchical modeling and identifiability. The framework is evaluated on two representative prior structures: Bayesian Linear Regression and a Hierarchical Gaussian Process. The results show that PLPFN's approximations remain close to asymptotically exact baselines, such as closed-form conjugate prior learning and Markov Chain Monte Carlo. Applying the learned priors to Bayesian Optimization shows that the PLPFN is competitive with state-of-the-art prior learning methods in terms of sample efficiency and final regret, both in- and out-of-distribution. The PLPFN framework is a valuable step in shifting prior specification in Bayesian inference from the complex manual translation of beliefs to specifying a set of related datasets.
...
Prior specification is a fundamental challenge in Bayesian inference. Traditionally, a prior represents a belief over the behavior of a statistical model based on expert knowledge. Such knowledge is not always available, or can be hard to express in a numerical form. As an answer, this thesis introduces the Prior-Learning Prior-Fitted Network (PLPFN), a transformer-based meta-learning model that amortizes prior learning. The model takes in a set of related tasks and outputs fully Bayesian uncertainty estimates over prior parameters in a single forward pass. The novelty lies in providing interpretable prior parameter outputs and examining the prior learning problem through the lens of hierarchical modeling and identifiability. The framework is evaluated on two representative prior structures: Bayesian Linear Regression and a Hierarchical Gaussian Process. The results show that PLPFN's approximations remain close to asymptotically exact baselines, such as closed-form conjugate prior learning and Markov Chain Monte Carlo. Applying the learned priors to Bayesian Optimization shows that the PLPFN is competitive with state-of-the-art prior learning methods in terms of sample efficiency and final regret, both in- and out-of-distribution. The PLPFN framework is a valuable step in shifting prior specification in Bayesian inference from the complex manual translation of beliefs to specifying a set of related datasets.
Bayesian Parameter Inference for Industrial Printer Models
Paving the Way for Probabilistic Nozzle Diagnostics
This thesis considers extensions of the standard independent hidden Markov model approach previously used by TNO for modelling printer nozzles. These extensions introduce parametrised transition probabilities and incorporate interactions between neighbouring nozzles to better capture the real-world printing process. The aim of this thesis is to investigate Bayesian methods to infer their model parameters. Although not the goal of this thesis, accurate parameter estimation paves the way for diagnosing nozzle malfunctions, ultimately improving printer performance.
The first method is a sampling algorithm with adaptive proposal distributions. We then present two variational inference (VI) algorithms, which aim to minimise the divergence between an approximate and true posterior of the unknown parameters. The first VI method makes use of model-specific approximations, while the latter is more flexible. The adaptive sampler, however, remains the most general of the three algorithms in addition to being asymptotically unbiased.
On synthetic data, all methods were capable of producing good estimates of the model parameters. When a good initialisation is available, the sampling method may be faster than the VI approaches; however, if no good start is known, the other methods may be preferred. The runtimes of the algorithms appeared to grow linearly with the number of nozzles. In addition, a linear scaling with the number of time steps appeared plausible.
While the methods performed well on the toy models studied here, their efficacy must be confirmed on more complex systems and demonstrated in real-world applications. If runtimes prove prohibitive, sub-sampling approaches or stochastic variational inference methods could be investigated. ...
The first method is a sampling algorithm with adaptive proposal distributions. We then present two variational inference (VI) algorithms, which aim to minimise the divergence between an approximate and true posterior of the unknown parameters. The first VI method makes use of model-specific approximations, while the latter is more flexible. The adaptive sampler, however, remains the most general of the three algorithms in addition to being asymptotically unbiased.
On synthetic data, all methods were capable of producing good estimates of the model parameters. When a good initialisation is available, the sampling method may be faster than the VI approaches; however, if no good start is known, the other methods may be preferred. The runtimes of the algorithms appeared to grow linearly with the number of nozzles. In addition, a linear scaling with the number of time steps appeared plausible.
While the methods performed well on the toy models studied here, their efficacy must be confirmed on more complex systems and demonstrated in real-world applications. If runtimes prove prohibitive, sub-sampling approaches or stochastic variational inference methods could be investigated. ...
This thesis considers extensions of the standard independent hidden Markov model approach previously used by TNO for modelling printer nozzles. These extensions introduce parametrised transition probabilities and incorporate interactions between neighbouring nozzles to better capture the real-world printing process. The aim of this thesis is to investigate Bayesian methods to infer their model parameters. Although not the goal of this thesis, accurate parameter estimation paves the way for diagnosing nozzle malfunctions, ultimately improving printer performance.
The first method is a sampling algorithm with adaptive proposal distributions. We then present two variational inference (VI) algorithms, which aim to minimise the divergence between an approximate and true posterior of the unknown parameters. The first VI method makes use of model-specific approximations, while the latter is more flexible. The adaptive sampler, however, remains the most general of the three algorithms in addition to being asymptotically unbiased.
On synthetic data, all methods were capable of producing good estimates of the model parameters. When a good initialisation is available, the sampling method may be faster than the VI approaches; however, if no good start is known, the other methods may be preferred. The runtimes of the algorithms appeared to grow linearly with the number of nozzles. In addition, a linear scaling with the number of time steps appeared plausible.
While the methods performed well on the toy models studied here, their efficacy must be confirmed on more complex systems and demonstrated in real-world applications. If runtimes prove prohibitive, sub-sampling approaches or stochastic variational inference methods could be investigated.
The first method is a sampling algorithm with adaptive proposal distributions. We then present two variational inference (VI) algorithms, which aim to minimise the divergence between an approximate and true posterior of the unknown parameters. The first VI method makes use of model-specific approximations, while the latter is more flexible. The adaptive sampler, however, remains the most general of the three algorithms in addition to being asymptotically unbiased.
On synthetic data, all methods were capable of producing good estimates of the model parameters. When a good initialisation is available, the sampling method may be faster than the VI approaches; however, if no good start is known, the other methods may be preferred. The runtimes of the algorithms appeared to grow linearly with the number of nozzles. In addition, a linear scaling with the number of time steps appeared plausible.
While the methods performed well on the toy models studied here, their efficacy must be confirmed on more complex systems and demonstrated in real-world applications. If runtimes prove prohibitive, sub-sampling approaches or stochastic variational inference methods could be investigated.
Statistical models often require more insights than the estimated values of a model. When these insights are not available by analytical means, one often resorts to resampling schemes. The most well known of which are bootstrapping and cross-validation. These techniques are very flexible and allow you to gain insight on whether you are overfitting your model, the distribution and bias of your estimators, the expected prediction error of regression models and many more. These then allow you to make statistical inferences such as hypothesis testing. The problem with these methods is the computational cost of repeatedly refitting models. When you are using Zestimators, the Swiss army infinitesimal jackknife is an approximation to what your parameters would be had they been estimated from a different weighting of the data. This approximation is very easy to compute and allows for very quick approximation of the reweighted parameter estimates. Giordano et al. 2020 provides a bound on the error of this approximation. This error bound relies on constants that are hard to compute, thus the result is not easily interpretable. In the case of applying the Swiss army infinitesimal jackknife to the case of estimating the parameter from a subset of the data, we build on the results by Giordano et al. 2020 by applying asymptotic theory. This leads to a much more interpretable result about the approximation error. By a simulation study we show that this asymptotic result still works well in a finite sample.
...
Statistical models often require more insights than the estimated values of a model. When these insights are not available by analytical means, one often resorts to resampling schemes. The most well known of which are bootstrapping and cross-validation. These techniques are very flexible and allow you to gain insight on whether you are overfitting your model, the distribution and bias of your estimators, the expected prediction error of regression models and many more. These then allow you to make statistical inferences such as hypothesis testing. The problem with these methods is the computational cost of repeatedly refitting models. When you are using Zestimators, the Swiss army infinitesimal jackknife is an approximation to what your parameters would be had they been estimated from a different weighting of the data. This approximation is very easy to compute and allows for very quick approximation of the reweighted parameter estimates. Giordano et al. 2020 provides a bound on the error of this approximation. This error bound relies on constants that are hard to compute, thus the result is not easily interpretable. In the case of applying the Swiss army infinitesimal jackknife to the case of estimating the parameter from a subset of the data, we build on the results by Giordano et al. 2020 by applying asymptotic theory. This leads to a much more interpretable result about the approximation error. By a simulation study we show that this asymptotic result still works well in a finite sample.
This thesis examines the connection between the Optimal Transport (OT) and the Schrödinger Bridge (SB) problem. Both analytical and numerical approaches to finding the optimal solution are discussed.
Originating from an engineering perspective, Optimal Transport seeks to find the minimal cost of transporting mass from one distribution to another. In the reformulation to the SB problem, these distributions are considered probability densities, and the solution can best be described as the most likely transition from the initial density to the other, expressed as a probability law. The optimal probability law is found by minimizing its Kullback-Leibler (KL) divergence when compared to a prior.
The theoretical mathematical foundations of the OT and SB problems are first explained. Here, the OT problem is approximated with an entropy regularisation term, which links it directly to the SB problem, by means of the KL-divergence.
To derive the optimal solution of the SB problem, its similarity with the discrete entropic OT problem is discussed. The existence of the unique solution to the discrete problem is found by its strict convexity. With the Lagrangian, the form of the solution it is established. In the continuous case, an iterative scheme to finding the solution is developed using similar computations. Scaling functions that define the unique optimal probability are obtained.
For one-dimensional Gaussian marginals, analytical update schemes are derived. First this is done for marginals with zero mean. For the Gaussians with non-zero mean, a more complicated derivation is explained, that considers the shifts in the mean.
Python implementations, found in appendices C, D, E and on GitHub 1, are made for all three update schemes, using numerical integration. The performance of the algorithms is compared, and their results show to be extremely similar. The general algorithm converges faster, but also takes more numerical approximations.
From the newly made computational general algorithm, an extension is derived that visualises intermediate marginals over a time grid, creating the "bridge". This is done for both Gaussian and non-Gaussian marginals. It can here be seen that generally, the algorithm performs better when marginals are "standard" Gaussian.
All in all, the Schrödinger Bridge provides a useful algorithm that can be used to solve potentially complicated OT problems. Both numerical implementations as analytical approaches give useful tools for finding solutions effectively. Additionally, bridge visualisations can be made neatly for a broad range of marginals.
Throughout the process of writing this thesis, the AI programme ChatGPTwas consulted for suggestions on academic sources with supporting arguments. Furthermore, it has been implemented to check for small errors in Python code as well as spelling, grammar and idiom mistakes, whilst maintaining the human-written text. ...
Originating from an engineering perspective, Optimal Transport seeks to find the minimal cost of transporting mass from one distribution to another. In the reformulation to the SB problem, these distributions are considered probability densities, and the solution can best be described as the most likely transition from the initial density to the other, expressed as a probability law. The optimal probability law is found by minimizing its Kullback-Leibler (KL) divergence when compared to a prior.
The theoretical mathematical foundations of the OT and SB problems are first explained. Here, the OT problem is approximated with an entropy regularisation term, which links it directly to the SB problem, by means of the KL-divergence.
To derive the optimal solution of the SB problem, its similarity with the discrete entropic OT problem is discussed. The existence of the unique solution to the discrete problem is found by its strict convexity. With the Lagrangian, the form of the solution it is established. In the continuous case, an iterative scheme to finding the solution is developed using similar computations. Scaling functions that define the unique optimal probability are obtained.
For one-dimensional Gaussian marginals, analytical update schemes are derived. First this is done for marginals with zero mean. For the Gaussians with non-zero mean, a more complicated derivation is explained, that considers the shifts in the mean.
Python implementations, found in appendices C, D, E and on GitHub 1, are made for all three update schemes, using numerical integration. The performance of the algorithms is compared, and their results show to be extremely similar. The general algorithm converges faster, but also takes more numerical approximations.
From the newly made computational general algorithm, an extension is derived that visualises intermediate marginals over a time grid, creating the "bridge". This is done for both Gaussian and non-Gaussian marginals. It can here be seen that generally, the algorithm performs better when marginals are "standard" Gaussian.
All in all, the Schrödinger Bridge provides a useful algorithm that can be used to solve potentially complicated OT problems. Both numerical implementations as analytical approaches give useful tools for finding solutions effectively. Additionally, bridge visualisations can be made neatly for a broad range of marginals.
Throughout the process of writing this thesis, the AI programme ChatGPTwas consulted for suggestions on academic sources with supporting arguments. Furthermore, it has been implemented to check for small errors in Python code as well as spelling, grammar and idiom mistakes, whilst maintaining the human-written text. ...
This thesis examines the connection between the Optimal Transport (OT) and the Schrödinger Bridge (SB) problem. Both analytical and numerical approaches to finding the optimal solution are discussed.
Originating from an engineering perspective, Optimal Transport seeks to find the minimal cost of transporting mass from one distribution to another. In the reformulation to the SB problem, these distributions are considered probability densities, and the solution can best be described as the most likely transition from the initial density to the other, expressed as a probability law. The optimal probability law is found by minimizing its Kullback-Leibler (KL) divergence when compared to a prior.
The theoretical mathematical foundations of the OT and SB problems are first explained. Here, the OT problem is approximated with an entropy regularisation term, which links it directly to the SB problem, by means of the KL-divergence.
To derive the optimal solution of the SB problem, its similarity with the discrete entropic OT problem is discussed. The existence of the unique solution to the discrete problem is found by its strict convexity. With the Lagrangian, the form of the solution it is established. In the continuous case, an iterative scheme to finding the solution is developed using similar computations. Scaling functions that define the unique optimal probability are obtained.
For one-dimensional Gaussian marginals, analytical update schemes are derived. First this is done for marginals with zero mean. For the Gaussians with non-zero mean, a more complicated derivation is explained, that considers the shifts in the mean.
Python implementations, found in appendices C, D, E and on GitHub 1, are made for all three update schemes, using numerical integration. The performance of the algorithms is compared, and their results show to be extremely similar. The general algorithm converges faster, but also takes more numerical approximations.
From the newly made computational general algorithm, an extension is derived that visualises intermediate marginals over a time grid, creating the "bridge". This is done for both Gaussian and non-Gaussian marginals. It can here be seen that generally, the algorithm performs better when marginals are "standard" Gaussian.
All in all, the Schrödinger Bridge provides a useful algorithm that can be used to solve potentially complicated OT problems. Both numerical implementations as analytical approaches give useful tools for finding solutions effectively. Additionally, bridge visualisations can be made neatly for a broad range of marginals.
Throughout the process of writing this thesis, the AI programme ChatGPTwas consulted for suggestions on academic sources with supporting arguments. Furthermore, it has been implemented to check for small errors in Python code as well as spelling, grammar and idiom mistakes, whilst maintaining the human-written text.
Originating from an engineering perspective, Optimal Transport seeks to find the minimal cost of transporting mass from one distribution to another. In the reformulation to the SB problem, these distributions are considered probability densities, and the solution can best be described as the most likely transition from the initial density to the other, expressed as a probability law. The optimal probability law is found by minimizing its Kullback-Leibler (KL) divergence when compared to a prior.
The theoretical mathematical foundations of the OT and SB problems are first explained. Here, the OT problem is approximated with an entropy regularisation term, which links it directly to the SB problem, by means of the KL-divergence.
To derive the optimal solution of the SB problem, its similarity with the discrete entropic OT problem is discussed. The existence of the unique solution to the discrete problem is found by its strict convexity. With the Lagrangian, the form of the solution it is established. In the continuous case, an iterative scheme to finding the solution is developed using similar computations. Scaling functions that define the unique optimal probability are obtained.
For one-dimensional Gaussian marginals, analytical update schemes are derived. First this is done for marginals with zero mean. For the Gaussians with non-zero mean, a more complicated derivation is explained, that considers the shifts in the mean.
Python implementations, found in appendices C, D, E and on GitHub 1, are made for all three update schemes, using numerical integration. The performance of the algorithms is compared, and their results show to be extremely similar. The general algorithm converges faster, but also takes more numerical approximations.
From the newly made computational general algorithm, an extension is derived that visualises intermediate marginals over a time grid, creating the "bridge". This is done for both Gaussian and non-Gaussian marginals. It can here be seen that generally, the algorithm performs better when marginals are "standard" Gaussian.
All in all, the Schrödinger Bridge provides a useful algorithm that can be used to solve potentially complicated OT problems. Both numerical implementations as analytical approaches give useful tools for finding solutions effectively. Additionally, bridge visualisations can be made neatly for a broad range of marginals.
Throughout the process of writing this thesis, the AI programme ChatGPTwas consulted for suggestions on academic sources with supporting arguments. Furthermore, it has been implemented to check for small errors in Python code as well as spelling, grammar and idiom mistakes, whilst maintaining the human-written text.
Conditioning Generative Diffusion Models
Training-free and Asymptotically Consistent
Generative diffusion is a machine learning technique to generate high-quality samples from complex data distributions. Much of its success can be attributed to the recently developed techniques that flexibly control the data generation process, without additional training effort. These methods control a pre-trained diffusion model towards specific regions of interest, which are determined by external information such as class labels, masks, or text descriptions. However, these approaches are typically based on heuristic guidance techniques and break the consistency on which the theoretical justification of generative diffusion relies. This is problematic when applying these controlled data generation techniques to tasks that are sensitive to distribution characteristics rather than the perceptual quality of individual samples. To this end, we introduce an asymptotically consistent approach for conditioning generative diffusion models without retraining the entire system. We use an importance sampling technique for simulating diffusion bridges, where multiple draws of a guided proposal process are reweighted to resemble paths of the true conditioned denoising process. A theoretical analysis shows that under certain assumptions, our approach has a vanishing error. In an empirical analysis, we find that specific nuances to the performance trade-off appear with a finite amount of computational effort. Specifically, the effectiveness of our approach highly depends on the choice of the proposal process and the allocation of computational effort towards independent runs of our algorithm.
...
Generative diffusion is a machine learning technique to generate high-quality samples from complex data distributions. Much of its success can be attributed to the recently developed techniques that flexibly control the data generation process, without additional training effort. These methods control a pre-trained diffusion model towards specific regions of interest, which are determined by external information such as class labels, masks, or text descriptions. However, these approaches are typically based on heuristic guidance techniques and break the consistency on which the theoretical justification of generative diffusion relies. This is problematic when applying these controlled data generation techniques to tasks that are sensitive to distribution characteristics rather than the perceptual quality of individual samples. To this end, we introduce an asymptotically consistent approach for conditioning generative diffusion models without retraining the entire system. We use an importance sampling technique for simulating diffusion bridges, where multiple draws of a guided proposal process are reweighted to resemble paths of the true conditioned denoising process. A theoretical analysis shows that under certain assumptions, our approach has a vanishing error. In an empirical analysis, we find that specific nuances to the performance trade-off appear with a finite amount of computational effort. Specifically, the effectiveness of our approach highly depends on the choice of the proposal process and the allocation of computational effort towards independent runs of our algorithm.
Variational inference comprises a family of statistical methods to obtain the optimal approximation of a target probability distribution using some reference class of distributions and a cost function, commonly the Kullback-Leibler (KL) divergence. Recent work on variational inference has yielded a fast, stable set of mean and covariance evolutions which dynamically yield variational Gaussian approximations via a restriction to Gaussian measures of the well-known JKO scheme. The sequence of Gaussian measures thus generated converges towards the KL-optimal Gaussian approximation of the VI target: it may also be used to approximate the entire sequence of distributions generated by a JKO gradient flow directed at this same target, thereby supporting practical usage of Gaussian VI as well as fast, approximate modelling of the Fokker-Planck PDE. However, it is not immediately clear whether this Gaussian sequence offers valid, helpful approximations of the original JKO gradient flow. In this work, three upper bounds for the sequence of Wasserstein-2 distances between the two gradient flows are obtained by exploiting the Riemannian structure of the W2 manifold and the shared properties of the Gaussian and JKO evolutions. Numerical simulations support the validity of these bounds and test their performance in both ordinary and exceptional scenarios. One of the bounds may be computed solely using the Gaussian evolution and the target potential, thus offering a tractable estimator for the suitability of variational Gaussian approximations which retains the attractive properties of Wasserstein distances whilst avoiding their computational demands.
...
Variational inference comprises a family of statistical methods to obtain the optimal approximation of a target probability distribution using some reference class of distributions and a cost function, commonly the Kullback-Leibler (KL) divergence. Recent work on variational inference has yielded a fast, stable set of mean and covariance evolutions which dynamically yield variational Gaussian approximations via a restriction to Gaussian measures of the well-known JKO scheme. The sequence of Gaussian measures thus generated converges towards the KL-optimal Gaussian approximation of the VI target: it may also be used to approximate the entire sequence of distributions generated by a JKO gradient flow directed at this same target, thereby supporting practical usage of Gaussian VI as well as fast, approximate modelling of the Fokker-Planck PDE. However, it is not immediately clear whether this Gaussian sequence offers valid, helpful approximations of the original JKO gradient flow. In this work, three upper bounds for the sequence of Wasserstein-2 distances between the two gradient flows are obtained by exploiting the Riemannian structure of the W2 manifold and the shared properties of the Gaussian and JKO evolutions. Numerical simulations support the validity of these bounds and test their performance in both ordinary and exceptional scenarios. One of the bounds may be computed solely using the Gaussian evolution and the target potential, thus offering a tractable estimator for the suitability of variational Gaussian approximations which retains the attractive properties of Wasserstein distances whilst avoiding their computational demands.
...
Master thesis
(2024)
-
T.M. Kamminga, Alexander Heinlein, M.B. van Gijzen, G.N.J.C. Bierkens, L. Bekker
Normalizing flows are a probabilistic method to estimate the underlying density of data samples. The method is flow based and non-parametric, with the aim of being flexible, but still computationally manageable. This report aims to explain the process of normalizing flows by the underlying principles of this probabilistic method. Both the method itself and the underlying principles are explained and reduced. An example of a normalizing flow is then implemented to showcase the method and support the underlying principles.
...
Normalizing flows are a probabilistic method to estimate the underlying density of data samples. The method is flow based and non-parametric, with the aim of being flexible, but still computationally manageable. This report aims to explain the process of normalizing flows by the underlying principles of this probabilistic method. Both the method itself and the underlying principles are explained and reduced. An example of a normalizing flow is then implemented to showcase the method and support the underlying principles.
A vital aspect of managing inflation risk is the use of inflation-indexed derivatives. Currently, inflation-indexed bonds and swaps are the primary instruments purchased by institutions. Inflation options (also known as inflation caps/floors) are also available in the market. Risk-neutral pricing of these derivatives is a difficult challenge due to the connection between inflation and interest rates.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published. ...
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published. ...
A vital aspect of managing inflation risk is the use of inflation-indexed derivatives. Currently, inflation-indexed bonds and swaps are the primary instruments purchased by institutions. Inflation options (also known as inflation caps/floors) are also available in the market. Risk-neutral pricing of these derivatives is a difficult challenge due to the connection between inflation and interest rates.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published.
In this thesis, the Heston model and its extensions to stochastic interest rates are investigated in the context of inflation-indexed derivatives. First, existing analytical pricing formulas and simulation methods are summarized. Then the multilevel Monte Carlo (MLMC) method is applied as a potent variance reduction technique. For the standard Heston model, the MLMC method reduces the computation costs by a factor of 10 to 50 for short maturities. The Python code implementing the applied methods is also published.
This thesis explores how forecasts of Dutch government bond yields can be improved by extending the current Dynamic Nelson-Siegel (DNS) model, used by the Dutch State Treasury Agency (DSTA), with stochastic volatility modeling and a Bayesian approach to parameter estimation and forecasting. The primary goal was to determine if the model extensions together with the Bayesian approach could improve the accuracy of yield forecasts given the highly volatile interest rate environment. In particular, we aimed to improve the "worst-case" forecasts, which we have defined as the upper bound of the 95% credible region with respect to the observed bond yields. To this end, we began with a baseline state-space model, resembling the current model in a state-space framework. Subsequently, we applied the findings from both in-sample and forecasting results as well as the findings from a literature review on volatility modeling to develop different models including two volatility models.
The volatility of the DNS model extensions is modeled as a GARCH process through the observation noise based on findings in the literature. This allowed for computationally efficient state estimation using a modified Kalman filter. Then, employing the Random Walk Metropolis algorithm for parameter estimation allowed us to use Bayesian multiple-step ahead forecasting. In particular, a comparative analysis of various models showed that while the current model performed better than expected, it was significantly outperformed in-sample by the DNS model with AR(1) observation noise (DNS-ARRW) and the DNS model with GARCH(1,1) observation noise volatility (DNS-OV). The Bayesian forecasting method particularly improved capturing the uncertainty of increasing yields in twelve-months ahead forecasts. Moreover, the two volatility models showed promising in-sample performance, but only one (DNS-OV) showed relatively good forecasting performance as well. Furthermore, the DNS-ARRW model consistently showed the best performance both in-sample and in forecasting.
In conclusion, the Bayesian approach to parameter estimation and forecasting proved effective in accounting for more variability in increasing forecast yields and simulating the direction of forecasts slightly better than the current MLE-based method. Moreover, the DNS-ARRW model showed significantly better worst-case forecasting performance, whereas the volatility models had a mixed performance. ...
The volatility of the DNS model extensions is modeled as a GARCH process through the observation noise based on findings in the literature. This allowed for computationally efficient state estimation using a modified Kalman filter. Then, employing the Random Walk Metropolis algorithm for parameter estimation allowed us to use Bayesian multiple-step ahead forecasting. In particular, a comparative analysis of various models showed that while the current model performed better than expected, it was significantly outperformed in-sample by the DNS model with AR(1) observation noise (DNS-ARRW) and the DNS model with GARCH(1,1) observation noise volatility (DNS-OV). The Bayesian forecasting method particularly improved capturing the uncertainty of increasing yields in twelve-months ahead forecasts. Moreover, the two volatility models showed promising in-sample performance, but only one (DNS-OV) showed relatively good forecasting performance as well. Furthermore, the DNS-ARRW model consistently showed the best performance both in-sample and in forecasting.
In conclusion, the Bayesian approach to parameter estimation and forecasting proved effective in accounting for more variability in increasing forecast yields and simulating the direction of forecasts slightly better than the current MLE-based method. Moreover, the DNS-ARRW model showed significantly better worst-case forecasting performance, whereas the volatility models had a mixed performance. ...
This thesis explores how forecasts of Dutch government bond yields can be improved by extending the current Dynamic Nelson-Siegel (DNS) model, used by the Dutch State Treasury Agency (DSTA), with stochastic volatility modeling and a Bayesian approach to parameter estimation and forecasting. The primary goal was to determine if the model extensions together with the Bayesian approach could improve the accuracy of yield forecasts given the highly volatile interest rate environment. In particular, we aimed to improve the "worst-case" forecasts, which we have defined as the upper bound of the 95% credible region with respect to the observed bond yields. To this end, we began with a baseline state-space model, resembling the current model in a state-space framework. Subsequently, we applied the findings from both in-sample and forecasting results as well as the findings from a literature review on volatility modeling to develop different models including two volatility models.
The volatility of the DNS model extensions is modeled as a GARCH process through the observation noise based on findings in the literature. This allowed for computationally efficient state estimation using a modified Kalman filter. Then, employing the Random Walk Metropolis algorithm for parameter estimation allowed us to use Bayesian multiple-step ahead forecasting. In particular, a comparative analysis of various models showed that while the current model performed better than expected, it was significantly outperformed in-sample by the DNS model with AR(1) observation noise (DNS-ARRW) and the DNS model with GARCH(1,1) observation noise volatility (DNS-OV). The Bayesian forecasting method particularly improved capturing the uncertainty of increasing yields in twelve-months ahead forecasts. Moreover, the two volatility models showed promising in-sample performance, but only one (DNS-OV) showed relatively good forecasting performance as well. Furthermore, the DNS-ARRW model consistently showed the best performance both in-sample and in forecasting.
In conclusion, the Bayesian approach to parameter estimation and forecasting proved effective in accounting for more variability in increasing forecast yields and simulating the direction of forecasts slightly better than the current MLE-based method. Moreover, the DNS-ARRW model showed significantly better worst-case forecasting performance, whereas the volatility models had a mixed performance.
The volatility of the DNS model extensions is modeled as a GARCH process through the observation noise based on findings in the literature. This allowed for computationally efficient state estimation using a modified Kalman filter. Then, employing the Random Walk Metropolis algorithm for parameter estimation allowed us to use Bayesian multiple-step ahead forecasting. In particular, a comparative analysis of various models showed that while the current model performed better than expected, it was significantly outperformed in-sample by the DNS model with AR(1) observation noise (DNS-ARRW) and the DNS model with GARCH(1,1) observation noise volatility (DNS-OV). The Bayesian forecasting method particularly improved capturing the uncertainty of increasing yields in twelve-months ahead forecasts. Moreover, the two volatility models showed promising in-sample performance, but only one (DNS-OV) showed relatively good forecasting performance as well. Furthermore, the DNS-ARRW model consistently showed the best performance both in-sample and in forecasting.
In conclusion, the Bayesian approach to parameter estimation and forecasting proved effective in accounting for more variability in increasing forecast yields and simulating the direction of forecasts slightly better than the current MLE-based method. Moreover, the DNS-ARRW model showed significantly better worst-case forecasting performance, whereas the volatility models had a mixed performance.
The widespread use of Markov Chain Monte Carlo (MCMC) methods for high-dimensional applications has motivated research into the scalability of these algorithms with respect to the dimension of the problem. Despite this, numerous problems concerning output analysis in high-dimensional settings have remained unaddressed. We present novel quantitative Gaussian approximation results for a broad range of MCMC algorithms. Notably, we analyse the dependency of the obtained approximation errors on the dimension of both the target distribution and the feature space. We demonstrate how these Gaussian approximations can be applied in output analysis. This includes determining the simulation effort required to guarantee Markov chain central limit theorems and consistent estimation of the variance and effective sample size in high-dimensional settings. We give quantitative convergence bounds for termination criteria and show that the termination time of a wide class of MCMC algorithms scales polynomially in dimension while ensuring a desired level of precision. Our results offer guidance to practitioners for obtaining appropriate standard errors and deciding the minimum simulation effort of MCMC algorithms in both multivariate and high-dimensional settings
...
The widespread use of Markov Chain Monte Carlo (MCMC) methods for high-dimensional applications has motivated research into the scalability of these algorithms with respect to the dimension of the problem. Despite this, numerous problems concerning output analysis in high-dimensional settings have remained unaddressed. We present novel quantitative Gaussian approximation results for a broad range of MCMC algorithms. Notably, we analyse the dependency of the obtained approximation errors on the dimension of both the target distribution and the feature space. We demonstrate how these Gaussian approximations can be applied in output analysis. This includes determining the simulation effort required to guarantee Markov chain central limit theorems and consistent estimation of the variance and effective sample size in high-dimensional settings. We give quantitative convergence bounds for termination criteria and show that the termination time of a wide class of MCMC algorithms scales polynomially in dimension while ensuring a desired level of precision. Our results offer guidance to practitioners for obtaining appropriate standard errors and deciding the minimum simulation effort of MCMC algorithms in both multivariate and high-dimensional settings
The computation of normalizing constants often brings (higher-dimensional) integrals, which could not be computed analytically or are computationally expensive. Stochastic methods will be explored to find an estimate for normalizing constants. These stochastic methods will be used to estimate the Bayes factor. With this Bayes factor can be searched for a best fitting model to data.
...
...
The computation of normalizing constants often brings (higher-dimensional) integrals, which could not be computed analytically or are computationally expensive. Stochastic methods will be explored to find an estimate for normalizing constants. These stochastic methods will be used to estimate the Bayes factor. With this Bayes factor can be searched for a best fitting model to data.
Since their introduction in 1993, particle filters are amongst the most popular algorithms for performing Bayesian inference on state space models that do not admit an analytical solution. In this thesis, we will present several particle filtering algorithms adapted to a class of models known as Piecewise Deterministic Markov Processes (PDMP), i.e. processes governed by one or more parameters that admit random jumps in their value at random times. Our work will focus on object tracking, the estimation of a target’s kinematic state over time from a sequence of noisy or incomplete measurements. Moreover, we will combine these techniques with Markov Chain Monte Carlo methods in order to infer the model parameters. We will perform sequential inference on both parameters and states by introducing an adaptation of the SMC2 to PDMPs. Finally, all algorithms will be tested both on simulated and real-world data (Piraeus AIS Dataset).
...
Since their introduction in 1993, particle filters are amongst the most popular algorithms for performing Bayesian inference on state space models that do not admit an analytical solution. In this thesis, we will present several particle filtering algorithms adapted to a class of models known as Piecewise Deterministic Markov Processes (PDMP), i.e. processes governed by one or more parameters that admit random jumps in their value at random times. Our work will focus on object tracking, the estimation of a target’s kinematic state over time from a sequence of noisy or incomplete measurements. Moreover, we will combine these techniques with Markov Chain Monte Carlo methods in order to infer the model parameters. We will perform sequential inference on both parameters and states by introducing an adaptation of the SMC2 to PDMPs. Finally, all algorithms will be tested both on simulated and real-world data (Piraeus AIS Dataset).
The spread of Covid-19 is modelled in this Bachelor thesis. This is done using two compartmental epidemiological models: the SIR-model and SIRS-model. These models divide the population into three groups, namely Susceptible, Infected and Recovered. Artificial data is generated based on the stochastic version of the models. The models are applied to this artificial data and to real data of a Covid-19 wave in Delft. Bayesian statistics are used to estimate the parameters of these models. This is done by using a Markov Chain Monte Carlo (MCMC) algorithm on the different data sets. The models are compared using the Root Mean Squared Error and the Bayes factor. For modelling the Covid-19 wave as it occurred in Delft, the SIRS-model is found to be preferred over the SIR-model. This preference is based on the Root Mean Squared Error (RMSE) and the Bayes Factor. Based on a MCMC simulation, the posterior mean value of the basic reproduction number is estimated to be 1.067. This value indicates a spread of the disease instead of it dying out.
...
The spread of Covid-19 is modelled in this Bachelor thesis. This is done using two compartmental epidemiological models: the SIR-model and SIRS-model. These models divide the population into three groups, namely Susceptible, Infected and Recovered. Artificial data is generated based on the stochastic version of the models. The models are applied to this artificial data and to real data of a Covid-19 wave in Delft. Bayesian statistics are used to estimate the parameters of these models. This is done by using a Markov Chain Monte Carlo (MCMC) algorithm on the different data sets. The models are compared using the Root Mean Squared Error and the Bayes factor. For modelling the Covid-19 wave as it occurred in Delft, the SIRS-model is found to be preferred over the SIR-model. This preference is based on the Root Mean Squared Error (RMSE) and the Bayes Factor. Based on a MCMC simulation, the posterior mean value of the basic reproduction number is estimated to be 1.067. This value indicates a spread of the disease instead of it dying out.
Recently, there has been an increase in literature about the Double Descent phenomenon for heavily over-parameterized models. Double Descent refers to the shape of the test risk curve, which can show a second descent in the over-parameterized regime, resulting in the remarkable combination of both low training and low test risk. However, much is still unknown about this behaviour. In this thesis we consider Double Descent and more specifically 'beneficial overfitting', meaning that the lowest test risk as a function of the number of parameters is achieved in the over-parameterized regime. We are mainly interested in under what conditions beneficial overfitting occurs. We start by exploring the test risk behaviour for simple linear regression models, with isotropic Gaussian, general Gaussian and sub-Gaussian covariates. For random feature selection and isotropic covariance, beneficial overfitting occurs for a large signal-to-noise ratio. For deterministic feature selection and isotropic covariance, beneficial overfitting occurs if we select features corresponding to the lowest weights. Without feature selection, beneficial overfitting occurs if the eigenvalues of the covariance matrix have a long, flat tail. In the second part of this thesis we check whether the same or similar results can be applied to other models as well. Specifically, we look at kernel regression, random Fourier features and a classification model. It seems that the linear regression results agree with the random Fourier features model, linear and quadratic kernel regression and classification model, but are not applicable for the Gaussian kernel regression case. Hence, more factors need to be considered, besides the eigenvalue behaviour of the covariance or kernel matrix and the way in which features are selected, to fully explain Double Descent and beneficial overfitting.
...
Recently, there has been an increase in literature about the Double Descent phenomenon for heavily over-parameterized models. Double Descent refers to the shape of the test risk curve, which can show a second descent in the over-parameterized regime, resulting in the remarkable combination of both low training and low test risk. However, much is still unknown about this behaviour. In this thesis we consider Double Descent and more specifically 'beneficial overfitting', meaning that the lowest test risk as a function of the number of parameters is achieved in the over-parameterized regime. We are mainly interested in under what conditions beneficial overfitting occurs. We start by exploring the test risk behaviour for simple linear regression models, with isotropic Gaussian, general Gaussian and sub-Gaussian covariates. For random feature selection and isotropic covariance, beneficial overfitting occurs for a large signal-to-noise ratio. For deterministic feature selection and isotropic covariance, beneficial overfitting occurs if we select features corresponding to the lowest weights. Without feature selection, beneficial overfitting occurs if the eigenvalues of the covariance matrix have a long, flat tail. In the second part of this thesis we check whether the same or similar results can be applied to other models as well. Specifically, we look at kernel regression, random Fourier features and a classification model. It seems that the linear regression results agree with the random Fourier features model, linear and quadratic kernel regression and classification model, but are not applicable for the Gaussian kernel regression case. Hence, more factors need to be considered, besides the eigenvalue behaviour of the covariance or kernel matrix and the way in which features are selected, to fully explain Double Descent and beneficial overfitting.
In this report, I investigate strategic decision making in the Formula E racing series for Porsche. Formula E is an electric car circuit racing series, where the main tasks of race strategy are allocating energy consumption across the race and timing mandatory "attack mode" activations (similar to small-scale pitstops). I work on a flawed existing tool that simulates Formula E races with the goal of identifying optimal strategies. By learning from Porsche's experience and comparing the tool against related literature, I investigate how to model a Formula E race.
To do this, I outline a method to measure the realism of our simulated races. By defining a race as a set of stochastic processes approximated with empirical data, I find a dissimilarity metric for the underlying probability distributions of two collections of races. My dissimilarity metric is based on Kantorovich's formulation of the optimal transport problem as applied to probability measures. The cost function used within this dissimilarity metric is based on how differently two drivers' individual races evolve from start to end, using gaps between times of arrival.
I introduce various alterations to the race model, based on experience and related literature. Through applying my metric to different versions of my race model, I evaluate these alterations and find improvements. Doing so shows measurable gains in the race model were achieved; this is also checked by measuring the prediction error of simulations versus the true observed race.
My approach yields a scalar metric that allows fast and straightforward tuning of any race model. It uses a high degree of aggregation to successfully compare otherwise not trivially comparable data. While already delivering a significantly improved model, my dissimilarity metric in conjunction with the prediction error-based metrics can help to find further improvements in the future. This is also independent of the used simulation method or format, since my metric only requires realisations of true and simulated races as input. ...
To do this, I outline a method to measure the realism of our simulated races. By defining a race as a set of stochastic processes approximated with empirical data, I find a dissimilarity metric for the underlying probability distributions of two collections of races. My dissimilarity metric is based on Kantorovich's formulation of the optimal transport problem as applied to probability measures. The cost function used within this dissimilarity metric is based on how differently two drivers' individual races evolve from start to end, using gaps between times of arrival.
I introduce various alterations to the race model, based on experience and related literature. Through applying my metric to different versions of my race model, I evaluate these alterations and find improvements. Doing so shows measurable gains in the race model were achieved; this is also checked by measuring the prediction error of simulations versus the true observed race.
My approach yields a scalar metric that allows fast and straightforward tuning of any race model. It uses a high degree of aggregation to successfully compare otherwise not trivially comparable data. While already delivering a significantly improved model, my dissimilarity metric in conjunction with the prediction error-based metrics can help to find further improvements in the future. This is also independent of the used simulation method or format, since my metric only requires realisations of true and simulated races as input. ...
In this report, I investigate strategic decision making in the Formula E racing series for Porsche. Formula E is an electric car circuit racing series, where the main tasks of race strategy are allocating energy consumption across the race and timing mandatory "attack mode" activations (similar to small-scale pitstops). I work on a flawed existing tool that simulates Formula E races with the goal of identifying optimal strategies. By learning from Porsche's experience and comparing the tool against related literature, I investigate how to model a Formula E race.
To do this, I outline a method to measure the realism of our simulated races. By defining a race as a set of stochastic processes approximated with empirical data, I find a dissimilarity metric for the underlying probability distributions of two collections of races. My dissimilarity metric is based on Kantorovich's formulation of the optimal transport problem as applied to probability measures. The cost function used within this dissimilarity metric is based on how differently two drivers' individual races evolve from start to end, using gaps between times of arrival.
I introduce various alterations to the race model, based on experience and related literature. Through applying my metric to different versions of my race model, I evaluate these alterations and find improvements. Doing so shows measurable gains in the race model were achieved; this is also checked by measuring the prediction error of simulations versus the true observed race.
My approach yields a scalar metric that allows fast and straightforward tuning of any race model. It uses a high degree of aggregation to successfully compare otherwise not trivially comparable data. While already delivering a significantly improved model, my dissimilarity metric in conjunction with the prediction error-based metrics can help to find further improvements in the future. This is also independent of the used simulation method or format, since my metric only requires realisations of true and simulated races as input.
To do this, I outline a method to measure the realism of our simulated races. By defining a race as a set of stochastic processes approximated with empirical data, I find a dissimilarity metric for the underlying probability distributions of two collections of races. My dissimilarity metric is based on Kantorovich's formulation of the optimal transport problem as applied to probability measures. The cost function used within this dissimilarity metric is based on how differently two drivers' individual races evolve from start to end, using gaps between times of arrival.
I introduce various alterations to the race model, based on experience and related literature. Through applying my metric to different versions of my race model, I evaluate these alterations and find improvements. Doing so shows measurable gains in the race model were achieved; this is also checked by measuring the prediction error of simulations versus the true observed race.
My approach yields a scalar metric that allows fast and straightforward tuning of any race model. It uses a high degree of aggregation to successfully compare otherwise not trivially comparable data. While already delivering a significantly improved model, my dissimilarity metric in conjunction with the prediction error-based metrics can help to find further improvements in the future. This is also independent of the used simulation method or format, since my metric only requires realisations of true and simulated races as input.