Prior Learning Through Transformers
V. Dakov (TU Delft - Electrical Engineering, Mathematics and Computer Science)
T.J. Viering – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
M.J.T. Reinders – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
G.N.J.C. Bierkens – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Prior specification is a fundamental challenge in Bayesian inference. Traditionally, a prior represents a belief over the behavior of a statistical model based on expert knowledge. Such knowledge is not always available, or can be hard to express in a numerical form. As an answer, this thesis introduces the Prior-Learning Prior-Fitted Network (PLPFN), a transformer-based meta-learning model that amortizes prior learning. The model takes in a set of related tasks and outputs fully Bayesian uncertainty estimates over prior parameters in a single forward pass. The novelty lies in providing interpretable prior parameter outputs and examining the prior learning problem through the lens of hierarchical modeling and identifiability. The framework is evaluated on two representative prior structures: Bayesian Linear Regression and a Hierarchical Gaussian Process. The results show that PLPFN's approximations remain close to asymptotically exact baselines, such as closed-form conjugate prior learning and Markov Chain Monte Carlo. Applying the learned priors to Bayesian Optimization shows that the PLPFN is competitive with state-of-the-art prior learning methods in terms of sample efficiency and final regret, both in- and out-of-distribution. The PLPFN framework is a valuable step in shifting prior specification in Bayesian inference from the complex manual translation of beliefs to specifying a set of related datasets.
Files
File under embargo until 22-06-2027