Can TabPFNs Learn Feature Importance?

Master Thesis (2026)
Author(s)

N. Annadanam (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

T.J. Viering – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

J.H. Krijthe – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

J. Yang – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
30-09-2026
Awarding Institution
Delft University of Technology
Programme
Computer Science, Data Science and Artificial Intelligence Technology
Faculty
Electrical Engineering, Mathematics and Computer Science
Page Views
9
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

In this work, we extend nanoTabPFN, a Prior-Data Fitted Network (PFN) for tabular classification to produce Shapley attributions alongside its predictive distribution. The output is a bucketed distribution over signed feature attribution values. The class prediction is derived as the additive sum of a baseline and per-feature contributions, ensuring the explanation is faithful to the prediction by construction. The attribution head is supervised by TreeSHAP values from a decision tree prior, connecting the explanation to the data generating process (DGP). The resulting attributions are interpreted as a posterior over Shapley values. The PFN’s distributional mechanism is transferred to feature attributions to yield calibrated explanation uncertainty. Across synthetic and real tabular benchmarks, the model recovers attributions with high fidelity (R2 ≈ 0.94-0.95 against exact Shapley on synthetic data), but at a small cost to accuracy. It produces calibrated attribution uncertainty, with negative log-likelihood (NLL) below the uninformative uniform baseline.

Files

License info not available