NN
N.H. Nanvani
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract both LiDAR segmentation with predictive confidence and semantic occupancy grids at arbitrary voxel resolutions. Rather than relying on domain-specific prompt engineering to counteract label flickering and temporal inconsistencies, SplatLabel directly distills continuous soft probabilities from 2D models, inherently resolving semantic ambiguities by aggregating predictions over time and space. To handle dynamic environments, we explicitly model the trajectories of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-computed 3D bounding boxes. Additionally, we guide scene geometry in unobserved regions by integrating 360-degree LiDAR priors via virtual depth maps. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.
...
While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract both LiDAR segmentation with predictive confidence and semantic occupancy grids at arbitrary voxel resolutions. Rather than relying on domain-specific prompt engineering to counteract label flickering and temporal inconsistencies, SplatLabel directly distills continuous soft probabilities from 2D models, inherently resolving semantic ambiguities by aggregating predictions over time and space. To handle dynamic environments, we explicitly model the trajectories of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-computed 3D bounding boxes. Additionally, we guide scene geometry in unobserved regions by integrating 360-degree LiDAR priors via virtual depth maps. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.