Label-Efficient 3D Pseudo-Label Refinement with Segmentation Foundation Model Priors

Master Thesis (2026)
Author(s)

P.A.R.M. Mignot (TU Delft - Mechanical Engineering)

Contributor(s)

Holger Caesar – Mentor (TU Delft - Mechanical Engineering)

S. Wang – Mentor (TU Delft - Mechanical Engineering)

Faculty
Mechanical Engineering
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
02-04-2026
Awarding Institution
Delft University of Technology
Programme
Mechanical Engineering, Vehicle Engineering, Cognitive Robotics
Faculty
Mechanical Engineering
Page Views
58
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

High-quality 3D annotations for LiDAR point clouds are expensive to obtain, while domain shift between source and target domains limits the direct use of pretrained 3D detectors. Auto-labeling offers a scalable alternative, but the quality of current pipelines remains largely bounded by LiDAR-based pseudo-label generation.

A key observation is that LiDAR auto-labeling pipelines can achieve high recall but produce many false positives, whereas image segmentation foundation models offer high-precision 2D detections with strong zero-shot generalisation. This suggests that the two modalities have complementary error profiles.

Building on this complementarity, we propose three modules that inject image segmentation priors into a LiDAR auto-labeling pipeline: rule-based confidence refinement, a weakly supervised learned refinement module, and mask-based 3D proposal generation.

On the View of Delft dataset, the learned module achieves the best pseudo-label quality among all configurations. Notably, geometric and confidence cues alone account for the majority of this improvement, with image features providing a consistent but modest additional gain. This indicates that the learned decision boundary, rather than the raw image signal, is the key enabler.

Files

License info not available