A.W. van der Vaart
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
10 records found
1
Causally interpretable meta-analysis (CIMA) extends conventional meta-analysis by transporting treatment effects from multiple randomized controlled trials to a prespecified target population. Existing CIMA methods assume that all shifted effect modifiers required for transportability are adequately captured by observed covariates. In practice, however, important effect modifiers may be partially observed or completely unobserved, limiting the applicability of these methods. In this thesis, we develop a proximal extension of CIMA for settings with unmeasured shifted effect modifiers. Building on proximal causal inference and proximal indirect comparison, we show that proximal indirect comparison cannot be directly adapted to CIMA because pooling multiple randomized trials induces an unmeasured treatment assignment mechanism in the source population. To address this challenge, we introduce proximal bridge functions and derive identification formulas for treatment-specific mean potential outcomes under a set of proximal assumptions. Based on these identification results, we propose outcome regression, inverse probability weighting, and doubly robust estimators for the target potential outcome means. Simulation studies demonstrate clear improvements over the original CIMA estimators in the presence of unmeasured shifted effect modifiers, while maintaining good finite-sample performance under a range of simulation settings.
...
Causally interpretable meta-analysis (CIMA) extends conventional meta-analysis by transporting treatment effects from multiple randomized controlled trials to a prespecified target population. Existing CIMA methods assume that all shifted effect modifiers required for transportability are adequately captured by observed covariates. In practice, however, important effect modifiers may be partially observed or completely unobserved, limiting the applicability of these methods. In this thesis, we develop a proximal extension of CIMA for settings with unmeasured shifted effect modifiers. Building on proximal causal inference and proximal indirect comparison, we show that proximal indirect comparison cannot be directly adapted to CIMA because pooling multiple randomized trials induces an unmeasured treatment assignment mechanism in the source population. To address this challenge, we introduce proximal bridge functions and derive identification formulas for treatment-specific mean potential outcomes under a set of proximal assumptions. Based on these identification results, we propose outcome regression, inverse probability weighting, and doubly robust estimators for the target potential outcome means. Simulation studies demonstrate clear improvements over the original CIMA estimators in the presence of unmeasured shifted effect modifiers, while maintaining good finite-sample performance under a range of simulation settings.
Stochastic partial differential equations (SPDEs) are mathematical models that describe the evolution of dynamical systems in space and time under the influence of random noise. The noise may represent the system’s inherent stochasticity, model uncertainties, or extrinsic stochastic forces acting on the system. By marrying the deterministic dynamics of classical partial differential equations with the influence of random noise, SPDEs allow for modeling of uncertainty in complex spatio-temporal phenomena in a wide range of fields, including geophysics, neuroscience, and finance.
In many real-world applications, a process modelled by an SPDE is observed only discretely and partially. As an example, consider an evolving firefront, whose fire intensity may be observed through satellite sensors at discrete points in time at a fixed spatial resolution. The SPDE model, jointly with the observed data, defines two statistical problems. Firstly, there is the estimation of the latent state of the system, given the observations. Typically, this is separated into two subproblems: the filtering problem, dealing with online state estimation as new observations become available, and the smoothing problem, dealing with offline state estimation given a full set of observations over a fixed time interval. Secondly, there is the question of parameter estimation. Typically, any SPDE model is governed by a set of parameters. These might, for example, represent physical constants that determine the dynamics of the system. In applications where such parameters are unknown, they are to be estimated jointly with the unobserved signal, based on the observed data.
This thesis develops Bayesian computational approaches to statistical inference for SPDEs. Both state estimation, in the online and online setting, as well as parameter estimation for discretely and partially observed semilinear SPDEs are addressed. It therefore fills a gap in the current literature on statistics for SPDEs, which has so far focused primarily on frequentist parameter estimation or state estimation of linear SPDEs. By introducing methodology that enables inference of the latent state, model calibration and uncertainty quantification for semilinear SPDEs, based on incomplete and noisy data, this work broadens the applicability of SPDEs in real-world settings.
In Chapter 2, a class of exponential measure transformations for SPDEs is introduced. Conditions are derived under which the transformed measure is of Girsanov-type such that the mild solution to the SPDE evolves according to yet another SPDE with an additional drift term. An application this result gives rise to the infinite-dimensional diffusion bridge - a mild solution to an SPDE conditioned on hitting a predefined terminal state. This generalises results previously known only for linear systems. Moreover, the guided process, the mild solution to a tractable SPDE that steers a process towards a terminal state, is introduced as an approximation to the diffusion bridge.
Chapter 3 develops sampling methodology for the intractable infinite-dimensional diffusion bridge. This serves as a fundamental building block for computational Bayesian approaches to inference for discretely observed SPDEs. In the main result of Chapter 3, conditions are derived under which absolute continuity holds between the laws of the guided process and the diffusion bridge. This legitimises the guided process as a proposal distribution for importance sampling or Metropolis-Hastings schemes that target the law of the infinite-dimensional diffusion bridge.
In Chapter 4, these ideas are extended to build methodology for all of the aforementioned statistical tasks. To address the filtering problem, a sequential Monte Carlo scheme is introduced, built upon the law of the guided process between observation times as a proposal distribution. The smoothing and parameter inference problems are solved by generalising the measure transformations of Chapter 2 to include conditioning and guiding based on multiple observations. Building on these transformations and a reparametrisation of the conditioned process, a Gibbs sampler is derived that samples from the joint posterior of the smoothed process and unknown model parameters.
...
In many real-world applications, a process modelled by an SPDE is observed only discretely and partially. As an example, consider an evolving firefront, whose fire intensity may be observed through satellite sensors at discrete points in time at a fixed spatial resolution. The SPDE model, jointly with the observed data, defines two statistical problems. Firstly, there is the estimation of the latent state of the system, given the observations. Typically, this is separated into two subproblems: the filtering problem, dealing with online state estimation as new observations become available, and the smoothing problem, dealing with offline state estimation given a full set of observations over a fixed time interval. Secondly, there is the question of parameter estimation. Typically, any SPDE model is governed by a set of parameters. These might, for example, represent physical constants that determine the dynamics of the system. In applications where such parameters are unknown, they are to be estimated jointly with the unobserved signal, based on the observed data.
This thesis develops Bayesian computational approaches to statistical inference for SPDEs. Both state estimation, in the online and online setting, as well as parameter estimation for discretely and partially observed semilinear SPDEs are addressed. It therefore fills a gap in the current literature on statistics for SPDEs, which has so far focused primarily on frequentist parameter estimation or state estimation of linear SPDEs. By introducing methodology that enables inference of the latent state, model calibration and uncertainty quantification for semilinear SPDEs, based on incomplete and noisy data, this work broadens the applicability of SPDEs in real-world settings.
In Chapter 2, a class of exponential measure transformations for SPDEs is introduced. Conditions are derived under which the transformed measure is of Girsanov-type such that the mild solution to the SPDE evolves according to yet another SPDE with an additional drift term. An application this result gives rise to the infinite-dimensional diffusion bridge - a mild solution to an SPDE conditioned on hitting a predefined terminal state. This generalises results previously known only for linear systems. Moreover, the guided process, the mild solution to a tractable SPDE that steers a process towards a terminal state, is introduced as an approximation to the diffusion bridge.
Chapter 3 develops sampling methodology for the intractable infinite-dimensional diffusion bridge. This serves as a fundamental building block for computational Bayesian approaches to inference for discretely observed SPDEs. In the main result of Chapter 3, conditions are derived under which absolute continuity holds between the laws of the guided process and the diffusion bridge. This legitimises the guided process as a proposal distribution for importance sampling or Metropolis-Hastings schemes that target the law of the infinite-dimensional diffusion bridge.
In Chapter 4, these ideas are extended to build methodology for all of the aforementioned statistical tasks. To address the filtering problem, a sequential Monte Carlo scheme is introduced, built upon the law of the guided process between observation times as a proposal distribution. The smoothing and parameter inference problems are solved by generalising the measure transformations of Chapter 2 to include conditioning and guiding based on multiple observations. Building on these transformations and a reparametrisation of the conditioned process, a Gibbs sampler is derived that samples from the joint posterior of the smoothed process and unknown model parameters.
...
Stochastic partial differential equations (SPDEs) are mathematical models that describe the evolution of dynamical systems in space and time under the influence of random noise. The noise may represent the system’s inherent stochasticity, model uncertainties, or extrinsic stochastic forces acting on the system. By marrying the deterministic dynamics of classical partial differential equations with the influence of random noise, SPDEs allow for modeling of uncertainty in complex spatio-temporal phenomena in a wide range of fields, including geophysics, neuroscience, and finance.
In many real-world applications, a process modelled by an SPDE is observed only discretely and partially. As an example, consider an evolving firefront, whose fire intensity may be observed through satellite sensors at discrete points in time at a fixed spatial resolution. The SPDE model, jointly with the observed data, defines two statistical problems. Firstly, there is the estimation of the latent state of the system, given the observations. Typically, this is separated into two subproblems: the filtering problem, dealing with online state estimation as new observations become available, and the smoothing problem, dealing with offline state estimation given a full set of observations over a fixed time interval. Secondly, there is the question of parameter estimation. Typically, any SPDE model is governed by a set of parameters. These might, for example, represent physical constants that determine the dynamics of the system. In applications where such parameters are unknown, they are to be estimated jointly with the unobserved signal, based on the observed data.
This thesis develops Bayesian computational approaches to statistical inference for SPDEs. Both state estimation, in the online and online setting, as well as parameter estimation for discretely and partially observed semilinear SPDEs are addressed. It therefore fills a gap in the current literature on statistics for SPDEs, which has so far focused primarily on frequentist parameter estimation or state estimation of linear SPDEs. By introducing methodology that enables inference of the latent state, model calibration and uncertainty quantification for semilinear SPDEs, based on incomplete and noisy data, this work broadens the applicability of SPDEs in real-world settings.
In Chapter 2, a class of exponential measure transformations for SPDEs is introduced. Conditions are derived under which the transformed measure is of Girsanov-type such that the mild solution to the SPDE evolves according to yet another SPDE with an additional drift term. An application this result gives rise to the infinite-dimensional diffusion bridge - a mild solution to an SPDE conditioned on hitting a predefined terminal state. This generalises results previously known only for linear systems. Moreover, the guided process, the mild solution to a tractable SPDE that steers a process towards a terminal state, is introduced as an approximation to the diffusion bridge.
Chapter 3 develops sampling methodology for the intractable infinite-dimensional diffusion bridge. This serves as a fundamental building block for computational Bayesian approaches to inference for discretely observed SPDEs. In the main result of Chapter 3, conditions are derived under which absolute continuity holds between the laws of the guided process and the diffusion bridge. This legitimises the guided process as a proposal distribution for importance sampling or Metropolis-Hastings schemes that target the law of the infinite-dimensional diffusion bridge.
In Chapter 4, these ideas are extended to build methodology for all of the aforementioned statistical tasks. To address the filtering problem, a sequential Monte Carlo scheme is introduced, built upon the law of the guided process between observation times as a proposal distribution. The smoothing and parameter inference problems are solved by generalising the measure transformations of Chapter 2 to include conditioning and guiding based on multiple observations. Building on these transformations and a reparametrisation of the conditioned process, a Gibbs sampler is derived that samples from the joint posterior of the smoothed process and unknown model parameters.
In many real-world applications, a process modelled by an SPDE is observed only discretely and partially. As an example, consider an evolving firefront, whose fire intensity may be observed through satellite sensors at discrete points in time at a fixed spatial resolution. The SPDE model, jointly with the observed data, defines two statistical problems. Firstly, there is the estimation of the latent state of the system, given the observations. Typically, this is separated into two subproblems: the filtering problem, dealing with online state estimation as new observations become available, and the smoothing problem, dealing with offline state estimation given a full set of observations over a fixed time interval. Secondly, there is the question of parameter estimation. Typically, any SPDE model is governed by a set of parameters. These might, for example, represent physical constants that determine the dynamics of the system. In applications where such parameters are unknown, they are to be estimated jointly with the unobserved signal, based on the observed data.
This thesis develops Bayesian computational approaches to statistical inference for SPDEs. Both state estimation, in the online and online setting, as well as parameter estimation for discretely and partially observed semilinear SPDEs are addressed. It therefore fills a gap in the current literature on statistics for SPDEs, which has so far focused primarily on frequentist parameter estimation or state estimation of linear SPDEs. By introducing methodology that enables inference of the latent state, model calibration and uncertainty quantification for semilinear SPDEs, based on incomplete and noisy data, this work broadens the applicability of SPDEs in real-world settings.
In Chapter 2, a class of exponential measure transformations for SPDEs is introduced. Conditions are derived under which the transformed measure is of Girsanov-type such that the mild solution to the SPDE evolves according to yet another SPDE with an additional drift term. An application this result gives rise to the infinite-dimensional diffusion bridge - a mild solution to an SPDE conditioned on hitting a predefined terminal state. This generalises results previously known only for linear systems. Moreover, the guided process, the mild solution to a tractable SPDE that steers a process towards a terminal state, is introduced as an approximation to the diffusion bridge.
Chapter 3 develops sampling methodology for the intractable infinite-dimensional diffusion bridge. This serves as a fundamental building block for computational Bayesian approaches to inference for discretely observed SPDEs. In the main result of Chapter 3, conditions are derived under which absolute continuity holds between the laws of the guided process and the diffusion bridge. This legitimises the guided process as a proposal distribution for importance sampling or Metropolis-Hastings schemes that target the law of the infinite-dimensional diffusion bridge.
In Chapter 4, these ideas are extended to build methodology for all of the aforementioned statistical tasks. To address the filtering problem, a sequential Monte Carlo scheme is introduced, built upon the law of the guided process between observation times as a proposal distribution. The smoothing and parameter inference problems are solved by generalising the measure transformations of Chapter 2 to include conditioning and guiding based on multiple observations. Building on these transformations and a reparametrisation of the conditioned process, a Gibbs sampler is derived that samples from the joint posterior of the smoothed process and unknown model parameters.
In proximal causal inference framework, the identification of average treatment effect (ATE) depends on finding the bridge functions. The bridge functions are functions about proxy variables used in the proximal standardization formulae. They are the solutions to two Fredholm integral equations of the first kind, whose existence is determined by Picard's conditions about the singular systems of two conditional expectation operators. However, since singular systems required by Picard's conditions are hard to determine, it is an extremely tough task to solve the bridge functions directly from the integral equations. Therefore, people turn to find estimators of the bridge functions. Many literatures have provided approaches to the estimators under certain assumptions although which inevitably restrict the feasibility of their application. In this thesis, we propose a kernel embedded estimator for the treatment confounding bridge function ($q$-bridge function) based on a dual kernel embedding method, under the assumption that there exist at least one bounded continuous $q$-bridge function for each treatment. In addition, we show the consistency of the $q$-bridge function estimator and give a consistent ATE estimator based on the proximal inverse probability weighted estimator.
...
In proximal causal inference framework, the identification of average treatment effect (ATE) depends on finding the bridge functions. The bridge functions are functions about proxy variables used in the proximal standardization formulae. They are the solutions to two Fredholm integral equations of the first kind, whose existence is determined by Picard's conditions about the singular systems of two conditional expectation operators. However, since singular systems required by Picard's conditions are hard to determine, it is an extremely tough task to solve the bridge functions directly from the integral equations. Therefore, people turn to find estimators of the bridge functions. Many literatures have provided approaches to the estimators under certain assumptions although which inevitably restrict the feasibility of their application. In this thesis, we propose a kernel embedded estimator for the treatment confounding bridge function ($q$-bridge function) based on a dual kernel embedding method, under the assumption that there exist at least one bounded continuous $q$-bridge function for each treatment. In addition, we show the consistency of the $q$-bridge function estimator and give a consistent ATE estimator based on the proximal inverse probability weighted estimator.
Chapter 1 (Caliper Matching): Caliper matching is used to estimate causal effects of a binary treatment from observational data by comparing matched treated and control units. Units are matched when their propensity scores, the conditional probability of receiving treatment given pretreatment covariates, are within a certain distance called caliper. So far, theoretical results on caliper matching are lacking, leaving practitioners with ad-hoc caliper choices and inference procedures. We bridge this gap by proposing a caliper that balances the quality and the number of matches. We prove that the resulting estimator of the average treatment effect, and average treatment effect on the treated, is asymptotically unbiased and normal at parametric rate. We describe the conditions under which semiparametric efficiency is obtainable, and show that when the parametric propensity score is estimated, the variance is increased for both estimands. Finally, we construct asymptotic confidence intervals for the two estimands.
Chapter 2 (Combining Experimental And Observational Data: The APOLLO Trial): In the APOLLO trial (Tol et al., 2022, 2024), we inferred the causal effects of two hip fracture treatments, Posterolateral Approach (PLA) and Direct Lateral Approach (DLA), on health outcomes of patients in the Netherlands. The starting point of the inference was a Randomised Experiment (RE), where patients were randomly assigned to PLA or DLA, independently of their baseline characteristics. In addition, data from a `Natural Experiment' (NE, or observational data) were also collected, under the plausible assumption that therein the allocation to PLA or DLA can be considered as good as random conditional on the patients' baseline characteristics. We estimated the average treatment effects of DLA versus PLA in the RE and the NE data separately, using a flexible and asymptotically efficient estimation strategy. We found no significant difference between PLA and DLA in any of the RE or the NE datasets. To improve the precision of the inference by increasing the sample size, we tested whether the RE and the NE datasets can be combined. Having found no evidence against combination, we estimated the average treatment effect on the combined dataset as well. Despite the improved precision, there was still no significant difference between PLA and DLA. Our conclusions were weakened by missing data, but they proved to be robust to our approach in handling the missingness and to our estimation strategy.
Chapter 3 (Private Double Robust Inference): Privacy mechanisms preserve the privacy of individuals in a sample by injecting noise into their sensitive data in a controlled manner, revealing only the noisy, privatised data to the statistician for inference purposes. The inference of a parameter exhibits a rate double robustness property when the large-sample bias of an estimator of the parameter is characterised by the product of the estimation errors of two other, auxiliary (or nuisance), often infinite-dimensional, parameters. We propose a novel class of rate double robust parameters whose novelty lies in the potentially nonlinear but smooth dependence on a low-dimensional regression parameter. Among others, this includes average treatment effects. We show that the properties of the sensitive-data model carry over to the privatised-data model by a suitable choice of the privacy mechanism, which, in general, means a total-variationally private mechanism. In particular, the double robustness property is retained, enabling efficient estimation from the privatised sample. We also find that the estimation of the nuisance parameters is not harder, albeit possibly less efficient and computationally more demanding, from the privatised sample compared to the sensitive sample for a given, suitable privacy mechanism. Indeed, if the estimation is feasible from the sensitive sample by some procedure, we can directly transport that procedure to the privatised setting. Lastly, we develop a private method of moments estimator for parametric models. This shows that in the private setting a parametric assumption about one nuisance parameter affords more flexible modelling and slower estimation of the other one, as is the well-known case in the nonprivate setting. ...
Chapter 2 (Combining Experimental And Observational Data: The APOLLO Trial): In the APOLLO trial (Tol et al., 2022, 2024), we inferred the causal effects of two hip fracture treatments, Posterolateral Approach (PLA) and Direct Lateral Approach (DLA), on health outcomes of patients in the Netherlands. The starting point of the inference was a Randomised Experiment (RE), where patients were randomly assigned to PLA or DLA, independently of their baseline characteristics. In addition, data from a `Natural Experiment' (NE, or observational data) were also collected, under the plausible assumption that therein the allocation to PLA or DLA can be considered as good as random conditional on the patients' baseline characteristics. We estimated the average treatment effects of DLA versus PLA in the RE and the NE data separately, using a flexible and asymptotically efficient estimation strategy. We found no significant difference between PLA and DLA in any of the RE or the NE datasets. To improve the precision of the inference by increasing the sample size, we tested whether the RE and the NE datasets can be combined. Having found no evidence against combination, we estimated the average treatment effect on the combined dataset as well. Despite the improved precision, there was still no significant difference between PLA and DLA. Our conclusions were weakened by missing data, but they proved to be robust to our approach in handling the missingness and to our estimation strategy.
Chapter 3 (Private Double Robust Inference): Privacy mechanisms preserve the privacy of individuals in a sample by injecting noise into their sensitive data in a controlled manner, revealing only the noisy, privatised data to the statistician for inference purposes. The inference of a parameter exhibits a rate double robustness property when the large-sample bias of an estimator of the parameter is characterised by the product of the estimation errors of two other, auxiliary (or nuisance), often infinite-dimensional, parameters. We propose a novel class of rate double robust parameters whose novelty lies in the potentially nonlinear but smooth dependence on a low-dimensional regression parameter. Among others, this includes average treatment effects. We show that the properties of the sensitive-data model carry over to the privatised-data model by a suitable choice of the privacy mechanism, which, in general, means a total-variationally private mechanism. In particular, the double robustness property is retained, enabling efficient estimation from the privatised sample. We also find that the estimation of the nuisance parameters is not harder, albeit possibly less efficient and computationally more demanding, from the privatised sample compared to the sensitive sample for a given, suitable privacy mechanism. Indeed, if the estimation is feasible from the sensitive sample by some procedure, we can directly transport that procedure to the privatised setting. Lastly, we develop a private method of moments estimator for parametric models. This shows that in the private setting a parametric assumption about one nuisance parameter affords more flexible modelling and slower estimation of the other one, as is the well-known case in the nonprivate setting. ...
Chapter 1 (Caliper Matching): Caliper matching is used to estimate causal effects of a binary treatment from observational data by comparing matched treated and control units. Units are matched when their propensity scores, the conditional probability of receiving treatment given pretreatment covariates, are within a certain distance called caliper. So far, theoretical results on caliper matching are lacking, leaving practitioners with ad-hoc caliper choices and inference procedures. We bridge this gap by proposing a caliper that balances the quality and the number of matches. We prove that the resulting estimator of the average treatment effect, and average treatment effect on the treated, is asymptotically unbiased and normal at parametric rate. We describe the conditions under which semiparametric efficiency is obtainable, and show that when the parametric propensity score is estimated, the variance is increased for both estimands. Finally, we construct asymptotic confidence intervals for the two estimands.
Chapter 2 (Combining Experimental And Observational Data: The APOLLO Trial): In the APOLLO trial (Tol et al., 2022, 2024), we inferred the causal effects of two hip fracture treatments, Posterolateral Approach (PLA) and Direct Lateral Approach (DLA), on health outcomes of patients in the Netherlands. The starting point of the inference was a Randomised Experiment (RE), where patients were randomly assigned to PLA or DLA, independently of their baseline characteristics. In addition, data from a `Natural Experiment' (NE, or observational data) were also collected, under the plausible assumption that therein the allocation to PLA or DLA can be considered as good as random conditional on the patients' baseline characteristics. We estimated the average treatment effects of DLA versus PLA in the RE and the NE data separately, using a flexible and asymptotically efficient estimation strategy. We found no significant difference between PLA and DLA in any of the RE or the NE datasets. To improve the precision of the inference by increasing the sample size, we tested whether the RE and the NE datasets can be combined. Having found no evidence against combination, we estimated the average treatment effect on the combined dataset as well. Despite the improved precision, there was still no significant difference between PLA and DLA. Our conclusions were weakened by missing data, but they proved to be robust to our approach in handling the missingness and to our estimation strategy.
Chapter 3 (Private Double Robust Inference): Privacy mechanisms preserve the privacy of individuals in a sample by injecting noise into their sensitive data in a controlled manner, revealing only the noisy, privatised data to the statistician for inference purposes. The inference of a parameter exhibits a rate double robustness property when the large-sample bias of an estimator of the parameter is characterised by the product of the estimation errors of two other, auxiliary (or nuisance), often infinite-dimensional, parameters. We propose a novel class of rate double robust parameters whose novelty lies in the potentially nonlinear but smooth dependence on a low-dimensional regression parameter. Among others, this includes average treatment effects. We show that the properties of the sensitive-data model carry over to the privatised-data model by a suitable choice of the privacy mechanism, which, in general, means a total-variationally private mechanism. In particular, the double robustness property is retained, enabling efficient estimation from the privatised sample. We also find that the estimation of the nuisance parameters is not harder, albeit possibly less efficient and computationally more demanding, from the privatised sample compared to the sensitive sample for a given, suitable privacy mechanism. Indeed, if the estimation is feasible from the sensitive sample by some procedure, we can directly transport that procedure to the privatised setting. Lastly, we develop a private method of moments estimator for parametric models. This shows that in the private setting a parametric assumption about one nuisance parameter affords more flexible modelling and slower estimation of the other one, as is the well-known case in the nonprivate setting.
Chapter 2 (Combining Experimental And Observational Data: The APOLLO Trial): In the APOLLO trial (Tol et al., 2022, 2024), we inferred the causal effects of two hip fracture treatments, Posterolateral Approach (PLA) and Direct Lateral Approach (DLA), on health outcomes of patients in the Netherlands. The starting point of the inference was a Randomised Experiment (RE), where patients were randomly assigned to PLA or DLA, independently of their baseline characteristics. In addition, data from a `Natural Experiment' (NE, or observational data) were also collected, under the plausible assumption that therein the allocation to PLA or DLA can be considered as good as random conditional on the patients' baseline characteristics. We estimated the average treatment effects of DLA versus PLA in the RE and the NE data separately, using a flexible and asymptotically efficient estimation strategy. We found no significant difference between PLA and DLA in any of the RE or the NE datasets. To improve the precision of the inference by increasing the sample size, we tested whether the RE and the NE datasets can be combined. Having found no evidence against combination, we estimated the average treatment effect on the combined dataset as well. Despite the improved precision, there was still no significant difference between PLA and DLA. Our conclusions were weakened by missing data, but they proved to be robust to our approach in handling the missingness and to our estimation strategy.
Chapter 3 (Private Double Robust Inference): Privacy mechanisms preserve the privacy of individuals in a sample by injecting noise into their sensitive data in a controlled manner, revealing only the noisy, privatised data to the statistician for inference purposes. The inference of a parameter exhibits a rate double robustness property when the large-sample bias of an estimator of the parameter is characterised by the product of the estimation errors of two other, auxiliary (or nuisance), often infinite-dimensional, parameters. We propose a novel class of rate double robust parameters whose novelty lies in the potentially nonlinear but smooth dependence on a low-dimensional regression parameter. Among others, this includes average treatment effects. We show that the properties of the sensitive-data model carry over to the privatised-data model by a suitable choice of the privacy mechanism, which, in general, means a total-variationally private mechanism. In particular, the double robustness property is retained, enabling efficient estimation from the privatised sample. We also find that the estimation of the nuisance parameters is not harder, albeit possibly less efficient and computationally more demanding, from the privatised sample compared to the sensitive sample for a given, suitable privacy mechanism. Indeed, if the estimation is feasible from the sensitive sample by some procedure, we can directly transport that procedure to the privatised setting. Lastly, we develop a private method of moments estimator for parametric models. This shows that in the private setting a parametric assumption about one nuisance parameter affords more flexible modelling and slower estimation of the other one, as is the well-known case in the nonprivate setting.
Proximal Causal Inference
Adjusting for the Unobserved
Causal relationships are at the heart of the scientific method. The causal revolution of the 21st century has opened the doors for many new approaches to quantify such relationships. In this thesis, we study the novel framework of proximal causal inference, which enables estimation of causal parameters even in the presence of unmeasured confounders, overcoming the limitations imposed by the Conditional Exchangeability assumption of the classic causal framework. In particular, we shall focus on determining and estimating the Causal Exposure Response Function (CERF) under this new set of assumptions. First, we introduce the problem and present a literature review of existing approaches to estimate bridge functions; then, we show theoretical results for extensions of classic linear results and a novel quasi- Bayesian method. This is then completed by showcasing performance on many simulated numerical examples and two real-world problems: Sustainable Causal Investing, and the effect of exercise on sleep.
...
Causal relationships are at the heart of the scientific method. The causal revolution of the 21st century has opened the doors for many new approaches to quantify such relationships. In this thesis, we study the novel framework of proximal causal inference, which enables estimation of causal parameters even in the presence of unmeasured confounders, overcoming the limitations imposed by the Conditional Exchangeability assumption of the classic causal framework. In particular, we shall focus on determining and estimating the Causal Exposure Response Function (CERF) under this new set of assumptions. First, we introduce the problem and present a literature review of existing approaches to estimate bridge functions; then, we show theoretical results for extensions of classic linear results and a novel quasi- Bayesian method. This is then completed by showcasing performance on many simulated numerical examples and two real-world problems: Sustainable Causal Investing, and the effect of exercise on sleep.
Causal inference with invalid instruments
Analysis of three different approaches for linear and non-linear models
Suppose that we want to infer the effect of a treatment on a certain outcome, where both the treatment and outcome are influenced by other variables. It has been well-established that in the linear setting, in case we know beforehand which of these other variables are instrumental (for the effect of the treatment on the outcome), we can infer the treatment effect in a consis-tent sense. This thesis analyses 3 methods that deals with the issue of unknown instrumental variables (IVs) and functional relationships in different ways to infer the treatment effect. The first method, Causal Inference with Invalid Instruments (CIII), assumes that we have a linear setting and a set with potential instrumental variables for whom a majority or plurality rule holds to obtain a robust confidence interval for the treatment effect. The second method, Anchor Regression (AR), only assumes a linear setting. By mediating between different meth-ods, the AR-estimator turns out the be robust to changes in the distribution of the sampled data. Lastly, Two Stage Curvature Identification (TSCI), does not require a linear setting or information on the IVs. Instead, it relies on the difference in functional form between the effect of the variables on the treatment and the effect of the variables on the outcome for consistent estimation and asymptotic normality. TSCI also provides a test for IV presence in the non-linear setting. In this thesis, I will explain the workings of these 3 methods, analyse their theoretical foundation and do simulation studies. Based on these analyses, I make several additions and suggestions to expand the theoretical scope and improve practical efficacy.3
...
...
Suppose that we want to infer the effect of a treatment on a certain outcome, where both the treatment and outcome are influenced by other variables. It has been well-established that in the linear setting, in case we know beforehand which of these other variables are instrumental (for the effect of the treatment on the outcome), we can infer the treatment effect in a consis-tent sense. This thesis analyses 3 methods that deals with the issue of unknown instrumental variables (IVs) and functional relationships in different ways to infer the treatment effect. The first method, Causal Inference with Invalid Instruments (CIII), assumes that we have a linear setting and a set with potential instrumental variables for whom a majority or plurality rule holds to obtain a robust confidence interval for the treatment effect. The second method, Anchor Regression (AR), only assumes a linear setting. By mediating between different meth-ods, the AR-estimator turns out the be robust to changes in the distribution of the sampled data. Lastly, Two Stage Curvature Identification (TSCI), does not require a linear setting or information on the IVs. Instead, it relies on the difference in functional form between the effect of the variables on the treatment and the effect of the variables on the outcome for consistent estimation and asymptotic normality. TSCI also provides a test for IV presence in the non-linear setting. In this thesis, I will explain the workings of these 3 methods, analyse their theoretical foundation and do simulation studies. Based on these analyses, I make several additions and suggestions to expand the theoretical scope and improve practical efficacy.3
Bayesian deep learning
Insights in the Bayesian paradigm for deep learning
In this thesis, we study a particle method for Bayesian deep learning. In particular, we look at the estimation of the parameters of an ensemble of Bayesian neural networks by means of this particle method, called Stein variational gradient descent (SVGD). This method iteratively updates a collection of parameters and it has the property that its update directions are chosen such that they optimally decrease the Kullback-Leibler divergence. We also study gradient flows of probability measures and show how gradient flows corresponding to functionals on the space of probability measures can induce particle flows. We formulate SVGD as a method in this space. In the regime of infinite particles we show results about convergence of SVGD. An existing convergence result for SVGD can be extended by showing that the probability measures, governing the collection of SVGD particles, are uniformly tight. We give conditions under which this holds.
...
In this thesis, we study a particle method for Bayesian deep learning. In particular, we look at the estimation of the parameters of an ensemble of Bayesian neural networks by means of this particle method, called Stein variational gradient descent (SVGD). This method iteratively updates a collection of parameters and it has the property that its update directions are chosen such that they optimally decrease the Kullback-Leibler divergence. We also study gradient flows of probability measures and show how gradient flows corresponding to functionals on the space of probability measures can induce particle flows. We formulate SVGD as a method in this space. In the regime of infinite particles we show results about convergence of SVGD. An existing convergence result for SVGD can be extended by showing that the probability measures, governing the collection of SVGD particles, are uniformly tight. We give conditions under which this holds.
Bayesian Sensitivity Analysis for a Missing Data Model
Incorporating Covariates via a Cox Model
In problems with missing data, the data are often considered to be missing at random. This assumption can not be checked from the data. We need to assess the sensitivity of study conclusions to violations of non-identifiable assumptions. This thesis performs Bayesian sensitivity analysis for a missing data model with life time outcomes and covariate information. The outcome distribution is modelled through a Cox model, with a beta process prior on the cumulative hazard function. We run experiments in a simulation study to test the performance of the model in scenarios with simulated data of several sample sizes. We show the validity of the model in the context of Bayesian sensitivity analysis, and propose extensions.
...
In problems with missing data, the data are often considered to be missing at random. This assumption can not be checked from the data. We need to assess the sensitivity of study conclusions to violations of non-identifiable assumptions. This thesis performs Bayesian sensitivity analysis for a missing data model with life time outcomes and covariate information. The outcome distribution is modelled through a Cox model, with a beta process prior on the cumulative hazard function. We run experiments in a simulation study to test the performance of the model in scenarios with simulated data of several sample sizes. We show the validity of the model in the context of Bayesian sensitivity analysis, and propose extensions.
When is subjective objective enough?
Frequentist analysis of Bayesian methods
In this thesis, we investigate the properties of Bayesian methods. In particular, we want to give frequentist guarantees for Bayesian methods. A Bayesian starts with specifying their apriori belief as a probability distribution, the prior distribution. The prior is their inherently subjective beliefs. After a Bayesian has specified their prior, they collect data and compute the posterior distribution. For a Bayesian, this posterior distribution encodes their new beliefs on the world. However, this prior was subjective. Thus the posterior is also subjective. So we can wonder, will this posterior distribution give a better representation of reality? Will it be more accurate? The posterior distribution quantifies a subjective belief of uncertainty. How reliable is this quantification of uncertainty?
These questions lie at the foundation of this thesis. They have been answered for certain classes of prior distributions. However, they have not been fully answered for all distributions in use. In this thesis, in the introduction, we explain the foundational statistical theory to study these questions. In particular, we show how to apply Schwartz theorem and the Bernstein-von Mises theorems to study posterior distributions. We then turn to novel research..... ...
These questions lie at the foundation of this thesis. They have been answered for certain classes of prior distributions. However, they have not been fully answered for all distributions in use. In this thesis, in the introduction, we explain the foundational statistical theory to study these questions. In particular, we show how to apply Schwartz theorem and the Bernstein-von Mises theorems to study posterior distributions. We then turn to novel research..... ...
In this thesis, we investigate the properties of Bayesian methods. In particular, we want to give frequentist guarantees for Bayesian methods. A Bayesian starts with specifying their apriori belief as a probability distribution, the prior distribution. The prior is their inherently subjective beliefs. After a Bayesian has specified their prior, they collect data and compute the posterior distribution. For a Bayesian, this posterior distribution encodes their new beliefs on the world. However, this prior was subjective. Thus the posterior is also subjective. So we can wonder, will this posterior distribution give a better representation of reality? Will it be more accurate? The posterior distribution quantifies a subjective belief of uncertainty. How reliable is this quantification of uncertainty?
These questions lie at the foundation of this thesis. They have been answered for certain classes of prior distributions. However, they have not been fully answered for all distributions in use. In this thesis, in the introduction, we explain the foundational statistical theory to study these questions. In particular, we show how to apply Schwartz theorem and the Bernstein-von Mises theorems to study posterior distributions. We then turn to novel research.....
These questions lie at the foundation of this thesis. They have been answered for certain classes of prior distributions. However, they have not been fully answered for all distributions in use. In this thesis, in the introduction, we explain the foundational statistical theory to study these questions. In particular, we show how to apply Schwartz theorem and the Bernstein-von Mises theorems to study posterior distributions. We then turn to novel research.....