KANs for Interpretable PDE Foundation Models
A.E. Mangos (TU Delft - Electrical Engineering, Mathematics and Computer Science)
D.M.J. Tax – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
J. Sun – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
M. Khosla – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Neural operator foundational models amortize the cost of repeatedly evaluating PDEs but their opaqueness and high dimensional inner representation makes them difficult to trust and understand, especially on out of distribution tasks. For certain (linear) PDE families this opacity is avoidable in principle since each instance is generated by a closed form (frequency domain) equation with only a handful of coefficients. As such, this thesis asks whether Kolmogorov-Arnold Networks (KANs) can be the building block that closes this accuracy and intepretability gap in 2 structurally different places: a drop-in component inside a Fourier Neural Operator (FNO) and as the foundation of a purpose built symbolic operator. We introduce KANSO (Kan Symbolic Operator), a compact KAN-first spectral operator whose architecture commits to a bilinear structure of constant coefficient 2D PDEs and exposes the equation learned directly (through the inherent intepretability of KANs), although a sparseness focused training regime is necessary to maintain intepretability with small accuracy degradation. Evaluating this 4 orders of magnitude smaller model than the black-box baseline FNO, it matches or exceeds it on zero-shot transfer throughout diffusion, transport and oscillatory families of PDEs. On the other hand, simply replacing a component of an FNO (the lift layer) with a KAN is only a conditional improvement, helping on smooth operators at large amount of parameters and degrades on oscillatory problems, which establishes that KANs are potential targeted improvements and not universal MLP replacements, motivating further the design of KANSO.