NLF-GS: A Mixed Mesh-Gaussian Representation for Generalizable and Drivable Human Avatars
J. He (TU Delft - Electrical Engineering, Mathematics and Computer Science)
C.A. Raman – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
P.S. Cesar Garcia – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
P. Kellnhofer – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Immersive telepresence requires avatar representations that preserve a person's identity and appearance while remaining responsive to changes in pose under practical deployment constraints such as low latency, compact transmission, and efficient rendering. Existing high-fidelity avatar methods often rely on person-specific optimization, while generalizable methods may struggle to maintain drivability under novel poses. To address this challenge, we present NLF-GS, a structured multi-view avatar reconstruction framework for learning generalizable and drivable mixed mesh-Gaussian avatars, inspired by Neural Localizer Fields (NLF). The core design of NLF-GS is to anchor Gaussian primitives to the faces of an SMPL-X body model and treat each Gaussian as a localized reconstruction query that predicts its geometry and appearance attributes in a feed-forward manner. This design combines the structural control of parametric human models with the rendering efficiency and local appearance flexibility of Gaussian primitives. Experiments show that NLF-GS achieves competitive reconstruction quality against representative generalizable baselines, while maintaining a compact representation and supporting novel-pose animation. The model generates new avatars at nearly 10 FPS and drives existing avatars at approximately 196 FPS on a RTX 4070 GPU. These results suggest that our mixed mesh-Gaussian representations provide a practical representation-level step toward scalable telepresence avatars.