Can TabPFNs Learn Feature Importance?
N. Annadanam (TU Delft - Electrical Engineering, Mathematics and Computer Science)
T.J. Viering – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
J.H. Krijthe – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
J. Yang – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
In this work, we extend nanoTabPFN, a Prior-Data Fitted Network (PFN) for tabular classification to produce Shapley attributions alongside its predictive distribution. The output is a bucketed distribution over signed feature attribution values. The class prediction is derived as the additive sum of a baseline and per-feature contributions, ensuring the explanation is faithful to the prediction by construction. The attribution head is supervised by TreeSHAP values from a decision tree prior, connecting the explanation to the data generating process (DGP). The resulting attributions are interpreted as a posterior over Shapley values. The PFN’s distributional mechanism is transferred to feature attributions to yield calibrated explanation uncertainty. Across synthetic and real tabular benchmarks, the model recovers attributions with high fidelity (R2 ≈ 0.94-0.95 against exact Shapley on synthetic data), but at a small cost to accuracy. It produces calibrated attribution uncertainty, with negative log-likelihood (NLL) below the uninformative uniform baseline.