A.W. van der Vaart
Please Note
18 records found
1
In causal inference, it is important to study the sensitivity of the conclusions to key assumptions. We perform sensitivity analysis of the assumption that missing outcomes are missing completely at random. We follow a Bayesian approach, which is nonparametric for the outcome distribution and can be combined with an informative prior on the sensitivity parameter. We give insight in the posterior and provide theoretical guarantees in the form of Bernstein-von Mises theorems for estimating the mean outcome. We study different parametrisations of the model involving Dirichlet process priors on the distribution of the outcome and on the distribution of the outcome conditional on the subject being treated. We show that these parametrisations incorporate a prior on the sensitivity parameter in different ways and discuss the relative merits. A key result on which the above theorems rely will be a general Bernstein-von Mises theorem for the normalised extended gamma process. We also present a simulation study, showing the performance of the methods in finite sample scenarios.
In this paper, we propose a novel Bayesian approach for nonparametric estimation in Wicksell’s problem. This has important applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in materials science to determine the 3D microstructure of a material, using its 2D cross-sections. We deviate from the classical Bayesian nonparametric approach, which would place a Dirichlet Process (DP) prior on the distribution function of the unobservables, by directly placing a DP prior on the distribution function of the observables. Our method offers computational simplicity due to the conjugacy of the posterior and allows for asymptotically efficient estimation by projecting the posterior onto the L2 subspace of increasing, right-continuous functions. Indeed, the resulting Isotonized Inverse Posterior (IIP) satisfies a Bernstein–von Mises (BvM) phenomenon with minimax asymptotic variance g0 (x)/2γ, where γ > 1/2 reflects the degree of Hölder continuity of the true cdf at x. Since the IIP gives automatic uncertainty quantification, it eliminates the need to estimate γ . Our results provide the first semiparametric Bernstein–von Mises theorem for projection-based posteriors with a DP prior in inverse problems.
Given a mild solution X to a semilinear stochastic partial differential equation (SPDE), we consider an exponential change of measure based on its infinitesimal generator L, defined in the topology of bounded pointwise convergence. The changed measure Ph depends on the choice of a function h in the domain of L. In our main result, we derive conditions on h for which the change of measure is of Girsanov-type. The process X under Ph is then shown to be a mild solution to another SPDE with an extra additive drift-term. We illustrate how different choices of h impact the law of X under Ph in selected applications. These include the derivation of an infinite-dimensional diffusion bridge as well as the introduction of guided processes for SPDEs, generalizing results known for finite-dimensional diffusion processes to the infinite-dimensional case.
We consider the accuracy of an approximate posterior distribution in nonparametric regression problems by combining posterior distributions computed on subsets of the data defined by the locations of the independent variables. We show that this approximate posterior retains the rate of recovery of the full data posterior distribution, where the rate of recovery adapts to the smoothness of the true regression function. As particular examples we consider Gaussian process priors based on integrated Brownian motion and the Matérn kernel augmented with a prior on the length scale. Besides theoretical guarantees we present a numerical study of the methods both on synthetic and real world data. We also propose a new aggregation technique, which numerically outperforms previous approaches. Finally, we demonstrate empirically that spatially distributed methods can adapt to local regularities, potentially outperforming the original Gaussian process. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Posterolateral or Direct Lateral Surgical Approach for Hemiarthroplasty After a Hip Fracture
A Randomized Clinical Trial Alongside a Natural Experiment
Importance: Hip fractures in older adults are serious injuries that result in disability, higher rates of illness and death, and a substantial strain on health care resources. High-quality evidence to improve hip fracture care regarding the surgical approach of hemiarthroplasty is lacking. Objective: To compare 6-month outcomes of the posterolateral approach (PLA) and direct lateral approach (DLA) for hemiarthroplasty in patients with acute femoral neck fracture. Design, Setting, and Participants: This multicenter, randomized clinical trial (RCT) comparing DLA and PLA was performed alongside a natural experiment (NE) at 14 centers in the Netherlands. Patients aged 18 years or older with an acute femoral neck fracture were included, with or without dementia. Secondary surgery of the hip, pathological fractures, or patients with multitrauma were excluded. Recruitment took place between February 2018 and January 2022. Treatment allocation was random or pseudorandom based on geographical location and surgeon preference. Statistical analysis was performed from July 2022 to September 2022. Exposure: Hemiarthroplasty using PLA or DLA. Main Outcome and Measures: The primary outcome was health-related quality of life 6 months after surgery, quantified with the EuroQol Group 5-Dimension questionnaire (EQ-5D-5L). Secondary outcomes included dislocations, fear of falling and falls, activities of daily living, pain, and reoperations. To improve generalizability, a novel technique was used for data fusion of the RCT and NE. Results: A total of 843 patients (542 [64.3%] female; mean [SD] age, 82.2 [7.5] years) participated, with 555 patients in the RCT (283 patients in the DLA group; 272 patients in the PLA group) and 288 patients in the NE (172 patients in the DLA group; 116 patients in the PLA group). In the RCT, mean EQ-5D-5L utility scores at 6 months were 0.50 (95% CI, 0.45-0.55) after DLA and 0.49 (95% CI, 0.44-0.54) after PLA, with 77% completeness. The between-group difference (-0.04 [95% CI, -0.11 to 0.04]) was not statistically significant nor clinically meaningful. Most secondary outcomes were comparable between groups, but PLA was associated with more dislocations than DLA (RCT: 15 of 272 patients [5.5%] in PLA vs 1 of 283 patients [0.4%] in DLA; NE: 6 of 113 patients [5.3%]) in PLA vs 2 of 175 patients [1.1%] in DLA). Data fusion resulted in an effect size of 0.00 (95% CI, -0.04 to 0.05) for the EQ-5D-5L and an odds ratio of 12.31 (95% CI, 2.77 to 54.70) for experiencing a dislocation after PLA. Conclusions and Relevance: This combined RCT and NE found that among patients treated with a cemented hemiarthroplasty after an acute femoral neck fracture, PLA was not associated with a better quality of life than DLA. Rates of dislocation and reoperation were higher after PLA. Randomized and pseudorandomized data yielded similar outcomes, which suggests a strengthening of these findings. Trial Registration: ClinicalTrials.gov Identifier: NCT04438226.
We obtain rates of contraction of posterior distributions in inverse problems with discrete observations. In a general setting of smoothness scales we derive abstract results for general priors, with contraction rates determined by discrete Galerkin approximation. The rate depends on the amount of prior concentration near the true function and the prior mass of functions with inferior Galerkin approximation. We apply the general result to non-conjugate series priors, showing that these priors give near optimal and adaptive recovery in some generality, Gaussian priors, and mixtures of Gaussian priors, where the latter are also shown to be near optimal and adaptive.
Combining test statistics from independent trials or experiments is a popular method of meta-analysis. However, there is very limited theoretical understanding of the power of the combined test, especially in high-dimensional models considering composite hypotheses tests. We derive a mathematical framework to study standard meta-analysis testing approaches in the context of the many normal means model, which serves as the platform to investigate more complex models. We introduce a natural and mild restriction on the meta-level combination functions of the local trials. This allows us to mathematically quantify the cost of compressing m trials into real-valued test statistics and combining these. We then derive minimax lower and matching upper bounds for the separation rates of standard combination methods for e.g. p-values and e-values, quantifying the loss relative to using the full, pooled data. We observe an elbow effect, revealing that in certain cases combining the locally optimal tests in each trial results in a sub-optimal meta-analysis method and develop approaches to achieve the global optima. We also explore the possible gains of allowing limited coordination between the trial designs. Our results connect meta-analysis with bandwidth constraint distributed inference and build on recent information theoretic developments in the latter field.
The Pitman-Yor process is a random probability distribution, that can be used as a prior distribution in a nonparametric Bayesian analy-sis. The process is of species sampling type and generates discrete distribu-tions, which yield of the order nσ different values (“species”) in a random sample of size n, ifthetypeσ is positive. Thus this type parameter can be set to target true distributions of various levels of discreteness, making the Pitman-Yor process an interesting prior in this case. It was previously shown that the resulting posterior distribution is consistent if and only if the true distribution of the data is discrete. In this paper we derive the dis-tributional limit of the posterior distribution, in the form of a (corrected) Bernstein-von Mises theorem, which previously was known only in the con-tinuous, inconsistent case. It turns out that the Pitman-Yor posterior distribution has good behaviour if the true distribution of the data is discrete with atoms that decrease not too slowly. Credible sets derived from the posterior distribution provide valid frequentist confidence sets in this case. For a general discrete distribution, the posterior distribution, although con-sistent, may contain a bias which does not converge to zero at the√n rate and invalidates posterior inference. We propose a bias correction that solves this problem. We also consider the effect of estimating the type parameter from the data, both by empirical Bayes and full Bayes methods. In a small simulation study we illustrate that without bias correction the coverage of credible sets can be arbitrarily low, also for some discrete distributions.
A common task in quality control is to determine a control limit for a product at the time of release that incorporates its risk of degradation over time. Such a limit for a given quality measurement will be based on empirical stability data, the intended shelf life of the product and the stability specification. The task is particularly important when the registered specifications for release and stability are equal. We discuss two relevant formulations and their implementations in both a frequentist and Bayesian framework. The first ensures that the risk of a batch failing the specification is comparable at release and at the end of shelf life. The second is to screen out batches at release time that are at high risk of failing the stability specification at the end of their shelf life. Although the second formulation seems more natural from a quality assurance perspective, it usually renders a control limit that is too stringent. In this paper we provide theoretical insight in this phenomenon, and introduce a heat-map visualisation that may help practitioners to assess the feasibility of implementing a limit under the second formulation. We also suggest a solution when infeasible. In addition, the current industrial benchmark is reviewed and contrasted to the two formulations. Computational algorithms for both formulations are laid out in detail, and illustrated on a dataset.
The features in a high-dimensional biomedical prediction problem are often well described by low-dimensional latent variables (or factors). We use this to include unlabeled features and additional information on the features when building a prediction model. Such additional feature information is often available in biomedical applications. Examples are annotation of genes, metabolites, or p-values from a previous study. We employ a Bayesian factor regression model that jointly models the features and the outcome using Gaussian latent variables. We fit the model using a computationally efficient variational Bayes method, which scales to high dimensions. We use the extra information to set up a prior model for the features in terms of hyperparameters, which are then estimated through empirical Bayes. The method is demonstrated in simulations and two applications. One application considers influenza vaccine efficacy prediction based on microarray data. The second application predicts oral cancer metastasis from RNAseq data.
Posterolateral or direct lateral approach for cemented hemiarthroplasty after femoral neck fracture (APOLLO)
Protocol for a multicenter randomized controlled trial with economic evaluation and natural experiment alongside
Background and purpose — The posterolateral and direct lateral surgical approach are the 2 most common surgical approaches for performing a hemiarthroplasty in patients with a hip fracture. It is unknown which surgical approach is preferable in terms of (cost-)effectiveness and quality of life. Methods and analysis — We designed a multicenter randomized controlled trial (RCT) with an economic evaluation and a natural experiment (NE) alongside. We will include 555 patients ≥ 18 years with an acute femoral neck fracture. The primary outcome is patient-reported health-related quality of life assessed with the EQ-5D-5L. Secondary outcomes include healthcare costs, complications, mortality, and balance (including fear of falling, actual falls, and injuries due to falling). An economic evaluation will be performed for quality adjusted life years (QALYs). We will use variable block randomization stratified for hospital. For continuous outcomes, we will use linear mixed-model analysis. Dichotomous secondary outcome measures will be analyzed using chi-square statistics and logistic regression models. Primary analyses are based on the intention-to-treat principle. Additional as treated analyses will be performed to evaluate the effect of protocol deviations. Study summary — (i) Largest RCT addressing the health-related patient outcome of the main surgical approaches of hemiarthroplasty. (ii) Focus on outcomes that are important for the patient. (iii) Pragmatic and inclusive RCT with few exclusion criteria, e.g., patients with dementia can participate. (iv) Natural experiment alongside to amplify the generalizability. (v) The first study conducting a costutility analysis comparing both surgical approaches.
Discrimination between potentially immunogenic protein aggregates and harmless pharmaceutical components, like silicone oil, is critical for drug development. Flow imaging techniques allow to measure and, in principle, classify subvisible particles in protein therapeutics. However, automated approaches for silicone oil discrimination are still lacking robustness in terms of accuracy and transferability. In this work, we present an image-based filter that can reliably identify silicone oil particles in protein therapeutics across a wide range of parenteral products. A two-step classification approach is designed for automated silicone oil droplet discrimination, based on particle images generated with a flow imaging instrument. Distinct from previously published methods, our novel image-based filter is trained using silicone oil droplet images only and is, thus, independent of the type of protein samples imaged. Benchmarked against alternative approaches, the proposed filter showed best overall performance in categorizing silicone oil and non-oil particles taken from a variety of protein solutions. Excellent accuracy was observed particularly for higher resolution images. The image-based filter can successfully distinguish silicone oil particles with high accuracy in protein solutions not used for creating the filter, showcasing its high transferability and potential for wide applicability in biopharmaceutical studies.
In Memoriam Kobus Oosterhoff (1933–2015)
Statistics as both a purely mathematical activity and an applied science
We consider nonparametric estimation of the Lévy measure of a hidden Lévy process driving a stationary Omstein-Uhlenbeck process which is observed at discrete time points. This Lévy measure can be expressed in terms of the canonical function of the stationary distribution of the Omstein-Uhlenbeck process, which is known to be self-decomposable. We propose an estimator for this canonical function based on a preliminary estimator of the characteristic function of the stationary distribution. We provide a suppport-reduction algorithm for the numerical computation of the estimator, and show that the estimator is asymptotically consistent under various sampling schemes. We also define a simple consistent estimator of the intensity parameter of the process. Along the way, a nonparametric procedure for estimating a self-decomposable density function is constructed, and it is shown that the Oenstein-Uhlenbeck process is β-mixing. Some general results on uniform convergence of random characteristic functions are included.