Circular Image

L. Nan

info

Please Note

48 records found

Accurate segmentation and analysis of individual trees from 3D point clouds is a crucial yet challenging task in urbanism and environmental studies. Most existing methods for tree instance segmentation suffer from either under- or over-segmentation errors, mainly due to the complex nature of the environments and the varying tree geometries. In this paper, we propose SATree, a novel structure-aware approach that directly identifies important tree structures, such as crowns and stems, from point clouds, enabling robust tree instance segmentation against tree overlaps and varying tree sizes. Our method leverages a multi-task learning framework that simultaneously performs (i) semantic segmentation to classify a point as crown, stem, or other; (ii) heatmap prediction to assign a heat value to each point based on 2D Gaussian kernels centered at tree stem locations; (iii) offset prediction to estimate point-wise offset vectors pointing to the instance centroid. Key to our approach is the stem localization module, where we fuse the semantic and heatmap predictions to reliably localize tree stems from the network outputs. After that, we utilize a graph-based shortest path algorithm to group individual tree points by integrating the learned offset embeddings. Extensive experiments on two public forestry datasets, TreeML and ForInstance, demonstrate that SATree consistently outperforms state-of-the-art methods in terms of AP, AP50, and AP25 scores, reducing significant under- or over-segmentation errors. Our research output supports downstream forestry inventory, 3D tree reconstruction, and fine-grained part segmentation of trees. Our source code of SATree is available at https://github.com/shenglandu/SATree. ...

Robust Multi-Modal 3D Multi-Object Tracking via Cross Correction

Journal article (2026) - Lipeng Gu, Xuefeng Yan, Weiming Wang, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei
Inaccurate detections remain a critical bottleneck in 3D multi-object tracking (MOT). Recent detection fusion-based methods incorporate camera detections as supplementary to reduce false detections and compensate for missing ones in LiDAR. However, their unidirectional camera-LiDAR correction lacks a feedback mechanism, precluding iterative mutual refinement between modalities for more robust LiDAR-based tracking. Inspired by the coarse-to-fine strategy in two-stage object detection, we introduce CrossTracker, a novel two-stage framework for online multi-modal 3D MOT. CrossTracker first constructs coarse camera and LiDAR trajectories independently, then performs trajectory fusion using both current and historical frames, without requiring future data. This ensures more robust mutual refinement between modalities. Specifically, CrossTracker comprises three core modules: i) the multi-modal modeling (M3) module, which fuses data from images, point clouds, and even planar geometry derived from images to establish a robust tracking constraint; ii) the coarse trajectory generation (C-TG) module, which independently generates coarse trajectories for both modalities using the M3 constraint; and iii) the trajectory fusion (TF) module, which applies mutual refinement between coarse LiDAR and camera trajectories through cross correction to ensure robust LiDAR trajectories. Extensive experiments show that CrossTracker outperforms 19 state-of-the-art methods, highlighting its effectiveness in leveraging the synergistic strengths of camera and LiDAR sensors for robust multi-modal 3D MOT. The code is available at https://github.com/lipeng-gu/CrossTracker. ...

From Multi-View Images to Text-Guided Neural Surface Edits

Conference paper (2026) - Nail Ibrahimli, Julian F.P. Kooij, Liangliang Nan
Implicit surface representations are valued for their compactness and continuity, but they pose significant challenges for editing. Despite recent advancements, existing methods often fail to preserve identity and maintain geometric consistency during editing. To address these challenges, we present NeuSEditor, a novel method for text-guided editing of neural implicit surfaces derived from multi-view images. NeuSEditor introduces an identity-preserving architecture that efficiently separates scenes into foreground and background, enabling precise modifications without altering the scene-specific elements. Our geometry-aware distillation loss significantly enhances rendering and geometric quality. Our method simplifies the editing workflow by eliminating the need for continuous dataset updates and source prompting. NeuSEditor outperforms recent state-of-the-art methods, delivering superior quantitative and qualitative results. for visual results, visit: neuseditor.github.io ...
Journal article (2026) - Tomislav Medic, Liangliang Nan
3D instance segmentation for laser scanning (LiDAR) point clouds remains a challenge in many remote sensing-related domains. Successful solutions typically rely on supervised deep learning and manual annotations, and consequently focus on objects that can be well delineated through visual inspection and manual labeling of point clouds. However, for tasks with more complex and cluttered scenes, such as in-field plant phenotyping in agriculture, such approaches are often infeasible. In this study, we tackle the task of in-field wheat head instance segmentation directly from terrestrial laser scanning (TLS) point clouds. To address the problem and circumvent the need for manual annotations, we propose a novel two-stage pipeline. To obtain the initial 3D instance proposals, the first stage uses 3D-to-2D multi-view projections, the Grounded SAM pipeline for zero-shot 2D object-centric segmentation, and multi-view label fusion. The second stage uses these initial proposals as noisy pseudo-labels to train a supervised 3D panoptic-style segmentation neural network. Our results demonstrate the feasibility of the proposed approach and show performance improvements relative to Wheat3DGS, a recent alternative solution for in-field wheat head instance segmentation without manual 3D annotations based on multi-view RGB images and 3D Gaussian Splatting, showcasing TLS as a competitive sensing alternative. Moreover, the results show that both stages of the proposed pipeline can deliver usable 3D instance segmentation without manual annotations, indicating promising, low-effort transferability to other comparable TLS-based point cloud segmentation tasks. ...

An efficient convex decomposition method of 3D building models for urban morphological analytics

Journal article (2026) - Yijie Wu, Fan Xue, Liangliang Nan, Longyong Wu, Jantien Stoter, Anthony G.O. Yeh
Urban morphological analytics on buildings is informative for sustainable development. 3D building massing features, such as courtyards and setbacks, reflect spatial organizations and circulations, while influence daylight access, ventilation, and shading. However, existing 3D GIS methods usually overlook such 3D massing features, further obscure morphological analytics and environmental assessment. This article proposes MorphCut, an efficient convex decomposition method that segments 3D shapes into mass-aligned parts. MorphCut leverages key morphological properties—planarity, regularity, and Gestalt laws—after a topological preprocessing step to enable mass-aware decomposition. Experiments on representative samples, ranging from small houses to complex skyscrapers, showed that MorphCut outperformed four baseline methods in (i) balancing convexity and compactness, (ii) aligning decomposed parts with building masses, and (iii) preserving geometric fidelity (average deviation: 0.25 m). An urban-scale validation on datasets from Delft and Hong Kong, comprising over 30,000 buildings across 18.3 km², demonstrated MorphCut’s robustness, scalability, and generalizability. MorphCut successfully decomposed 98% of buildings in low-rise regions (+78% over the second-best method) and 93% in high-rise areas (+2%), completing processing in 13 hours (3 hours faster). These results position MorphCut as a foundational 3D GIS tool for large-scale, mass-aware morphological analysis, with implications for digital twins, sustainable planning, and environmental modeling. ...
Journal article (2026) - Jiaxiu Zhang, Wei Zhao, Ran Chen, Liangliang Nan, Wenhao Chen, Mingqiang Wei
Porous sandwich structures, particularly aluminum foam sandwiches (AFS), are widely used in lightweight and impact-resistant applications, yet their mechanical performance remains difficult to predict due to irregular and multiscale pore morphologies. Traditional constitutive models and current deep learning methods fall short in capturing the complex structure–property relationships of those materials. Accordingly, this work proposes a three-dimensional (3D) pore cloud representation learning method tailored for energy absorption prediction. A novel digital descriptor, termed the pore cloud, is constructed from 3D scans of real AFS cores to preserve detailed pore-level geometric and topological information. A comprehensive structure–property dataset is subsequently generated by integrating these pore cloud features with energy absorption data obtained through finite element analysis (FEA). Furthermore, this work develops PoreNet, a point cloud-based deep learning architecture that learns the direct mapping from mesoscale pore morphology to macroscopic mechanical response. The experimental results demonstrate that PoreNet achieves a high prediction accuracy of 95.12%, robust generalization across variable porosities, and fast convergence within 30 min on a consumer-grade Graphics Processing Unit (GPU). It outperforms both traditional analytical models and baseline neural networks. In addition, this study demonstrates the effectiveness of pore-level geometric learning in structure–property modeling and offers a scalable, data-driven framework for the design and optimization of advanced porous sandwich composites. The dataset and the proposed algorithm are publicly available at https://crescentrosexx.github.io/pore-net/. ...

Explore Open-Domain Image-to-Point Cloud Registration Using Topology Relationship

Conference paper (2025) - Pei An, Jiaqi Yang, Muyao Peng, You Yang, Qiong Liu, Jie Ma, Liangliang Nan
Image-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationship perspective. Firstly, we find that topology relationship reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationship. After that, to construct and leverage the topology relationship between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. ...
Conference paper (2025) - Y. Lin, S. Wang, L. Nan, J.F.P. Kooij, Holger Caesar
Scene flow estimation aims to recover per-point motion from two adjacent LiDAR scans. However, in real-world applications such as autonomous driving, points rarely move independently of others, especially for nearby points belonging to the same object, which often share the same motion. Incorporating this locally rigid motion constraint has been a key challenge in self-supervised scene flow estimation, which is often addressed by post-processing or appending extra regularization. While these approaches are able to improve the rigidity of predicted flows, they lack an architectural inductive bias for local rigidity within the model structure, leading to suboptimal learning efficiency and inferior performance. In contrast, we enforce local rigidity with a lightweight add-on module in neural network design, enabling end-to-end learning. We design a discretized voting space that accommodates all possible translations and then identify the one shared by nearby points by differentiable voting. Additionally, to ensure computational efficiency, we operate on pillars rather than points and learn representative features for voting per pillar. We plug the Voting Module into popular model designs and evaluate its benefit on Argoverse 2 and Waymo datasets. We outperform baseline works with only marginal compute overhead. Code is available at https://github.com/tudelft-iv/VoteFlow. ...
Journal article (2025) - Pei An, You Yang, Jiaqi Yang, Muyao Peng, Qiong Liu, Liangliang Nan
Image-to-point-cloud (I2P) registration is a fundamental yet challenging problem in computer vision. Despite significant advances in deep learning, I2P registration struggles with correspondence accuracy when training samples are limited. To address this challenge, we propose a Beltrami flow based I2P registration method termed Flow-I2P. From the perspective of information geometry, I2P registration can be reframed as a manifold alignment problem. Our in-depth analysis shows that Beltrami flow enhances I2P registration by improving manifold alignment quality. Building on this analysis, we introduce a Beltrami flow based cross-modality feature interaction layer, B-flow, to progressively refine manifold alignment. To reduce memory and computation demands, B-flow is then optimized into C-flow through the incorporation of feature covariance-based attention. We further enhance I2P registration performance by developing Flow-I2P, which incorporates normal features, stacked C-flow layers, and a two-stage training strategy. To evaluate the registration performance of Flow-I2P, we conduct extensive experiments on five indoor and outdoor datasets, including RGB-D V2, 7-Scenes, ScanNet, KITTI, and a self-collected dataset. Our results indicate that Flow-I2P achieves higher inlier ratio (IR) and registration recall (RR) compared to state-of-the-art methods. We conclude that Flow-I2P significantly enhances I2P registration with superior capabilities. ...

Quantitative Leaf Reconstruction From TLS Point Clouds

Journal article (2025) - Guangpeng Fan, Liangliang Xu, Jiani Guo, Ruoyoulan Wang, Haoran Zhao, Hao Lu, Jinhu Wang, Di Wang, Feixiang Chen, Liangliang Nan
Quantitatively reconstructing the 3-D structure of individual leaves within tree canopies is critical for understanding forest function and environmental responses to climate change. While quantitative structure models (QSMs) using terrestrial laser scanning (TLS) effectively capture woody structures, they lack the capability to accurately reconstruct nonwoody leaf components. This study proposes accurate and detailed leaf (AdLeaf), a novel approach for fine-scale reconstruction of individual leaves using TLS point clouds. AdLeaf combines wood–leaf separation, individual leaf segmentation, detection and repair of incomplete leaves, explicit reconstruction, and parameter extraction. It automates semantic segmentation at the tree scale to separate woody and leafy components. Instance segmentation is refined through similarity graphs. Incomplete leaves are detected and repaired using shape concavity analysis and symmetry-based mirroring. AdLeaf enables direct measurement of leaf attributes, including count, area, inclination, volume, and azimuth. Validation using field scans, synthetic data, and both in situ and destructive measurements shows high accuracy: leaf counting errors ranged from 0.58% to 8.23% for trees with 201–4000 leaves. Reconstructed leaf geometries had mean and standard deviations (SDs) below 0.83 and 0.70 cm, respectively. Leaf area measurements (10–180 cm2) achieved a coefficient of determination (R2) of 0.95, a bias of −0.20 cm2, and a root-mean-square error of 5.63 cm2. Incomplete leaf detection errors were below 28%, with the repaired area relative root-mean-square error (rRMSE) reduced by 9.4%. By addressing QSM limitations, AdLeaf enables explicit 3-D leaf reconstructions that support detailed analysis of canopy light interception, spatial heterogeneity, and photosynthesis. It provides a robust framework for linking leaf structure to function at the tree level, advancing forest structure and radiative transfer research. ...

Self-supervised Point Cloud Learning via Joint Completion and Generation

Journal article (2025) - Yun Liu, Peng Li, Xuefeng Yan, Liangliang Nan, Bing Wang, Honghua Chen, Lina Gong, Wei Zhao, Mingqiang Wei
The core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points’ representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance. ...

Multi-View Consistent Artistic Style Transfer

Conference paper (2024) - N. Ibrahimli, J.F.P. Kooij, L. Nan
We introduce MuVieCAST, a modular multi-view consistent style transfer network architecture that enables consistent style transfer between multiple viewpoints of the same scene. This network architecture supports both sparse and dense views, making it versatile enough to handle a wide range of multi-view image datasets. The approach consists of three modules that perform specific tasks related to style transfer, namely content preservation, image transformation, and multi-view consistency enforcement. We extensively evaluate our approach across multiple application domains including depth-map-based point cloud fusion, mesh reconstruction, and novel-view synthesis. Our experiments reveal that the proposed framework achieves an exceptional generation of stylized images, exhibiting consistent outcomes across perspectives. A user study focusing on novel-view synthesis further confirms these results, with approximately 68% of cases participants expressing a preference for our generated outputs compared to the recent state-of-the-art method. Our modular framework is extensible and can easily be integrated with various backbone architectures, making it a flexible solution for multi-view style transfer. ...
Journal article (2024) - Nima Forouzandeh, Eleonora Brembilla, Liangliang Nan, Jantien Stoter, Alstan Jakubiec
Optimizing the built environment via simulations of building models hinges on standardizing data acquisition. In this research, we put forward distinct levels of detail for geometry and material inputs, specifically tailored for indoor daylight applications. We primarily focus on understanding the uncertainties arising from imprecise estimations of material optical properties and incomplete geometrical inputs in climate-based indoor daylight simulations. Employing a Monte Carlo approach, we analyzed six office and teaching spaces, creating 20 variations for each by altering geometrical completeness and material accuracy. The technique of excluding non-permanent objects below certain sizes in four graduated steps was used to derive and test the impact of various geometrical levels of detail. Our findings reveal that different levels of geometrical completeness lead to errors ranging from 1.08% to 18.05%. Additionally, a twofold increase in simulation time was noted when geometrical detail was enhanced relative to the most basic model. Errors stemming from imprecise definitions of material optical properties showed a normal distribution. The uncertainty in simulation outcomes showed a linear rise with increasing input material uncertainty, lying between 10% to 30%, depending on space configurations. We observed heightened uncertainty near openings, attributed to window transmittance effects. The research underscores that daylight predictions are markedly more sensitive to transmittance uncertainties than to those in reflectance, regardless of the window-to-floor ratio. These insights may help to guide a more efficient data acquisition process of indoor spaces for daylight simulations. ...
Journal article (2024) - Minglei Li, Shu Peng, Liangliang Nan
We propose a concept of hybrid geometry sets for registering cross-source geometric data. Specifically, our method focuses on the coarse registration of geometric data obtained from laser scanning and photogrammetric reconstruction. Due to different characteristics (e.g., variations in noise levels, density, and scales), achieving accurate registration between these data becomes a challenging task. The proposed method uses geometric structures to construct hybrid geometry sets, and the geometric relations between the elements of a hybrid geometry set are encoded in a hybrid feature space. This enables effective and efficient similarity query and correspondence establishment between the hybrid geometry sets. The proposed global registration method works in three steps. Firstly, a set of hybrid geometry sets is constructed using extracted planes and intersection lines. Then the features of the hybrid geometry sets are computed to encode the relative pose and topological relationships between the extracted planes and intersection lines, and their correspondences between the two inputs are established by querying hybrid geometry sets with similar features. Finally, the global registration parameters are calculated using the correspondences, and the registration result is further refined through continuous optimization. The robustness of the method has been evaluated using different real-world cross-source geometric data of urban scenes. Extensive comparisons with state-of-the-art algorithms have also demonstrated its effectiveness. ...

A lightweight framework for effective and efficient point cloud analysis

Journal article (2024) - Lipeng Gu, Xuefeng Yan, Liangliang Nan, Dingkun Zhu, Honghua Chen, Weiming Wang, Mingqiang Wei
The conventional wisdom in point cloud analysis predominantly explores 3D geometries. It is often achieved through the introduction of intricate learnable geometric extractors in the encoder or by deepening networks with repeated blocks. However, these methods contain a significant number of learnable parameters, resulting in substantial computational costs and imposing memory burdens on CPU/GPU. Moreover, they are primarily tailored for object-level point cloud classification and segmentation tasks, with limited extensions to crucial scene-level applications, such as autonomous driving. To this end, we introduce PointeNet, an efficient network designed specifically for point cloud analysis. PointeNet distinguishes itself with its lightweight architecture, low training cost, and plug-and-play capability, while also effectively capturing representative features. The network consists of a Multivariate Geometric Encoding (MGE) module and an optional Distance-aware Semantic Enhancement (DSE) module. MGE employs operations of sampling, grouping, pooling, and multivariate geometric aggregation to lightweightly capture and adaptively aggregate multivariate geometric features, providing a comprehensive depiction of 3D geometries. DSE, designed for real-world autonomous driving scenarios, enhances the semantic perception of point clouds, particularly for distant points. Our method demonstrates flexibility by seamlessly integrating with a classification/segmentation head or embedding into off-the-shelf 3D object detection networks, achieving notable performance improvements at a minimal cost. Extensive experiments on object-level datasets, including ModelNet40, ScanObjectNN, ShapeNetPart, and the scene-level dataset KITTI, demonstrate the superior performance of PointeNet over state-of-the-art methods in point cloud analysis. Notably, PointeNet outperforms PointMLP with significantly fewer parameters on ModelNet40, ScanObjectNN, and ShapeNetPart, and achieves a substantial improvement of over 2% in 3DAPR40 for PointRCNN on KITTI with a minimal parameter cost of 1.4 million. Code is publicly available at https://github.com/lipeng-gu/PointeNet ...
Journal article (2024) - Xiaoxin Mi, Zhen Dong, Zhipeng Cao, Bisheng Yang, Zhen Cao, Chao Zheng, Jantien Stoter, Liangliang Nan
Accurate lane maps with semantics are crucial for various applications, such as high-definition maps (HD Maps), intelligent transportation systems (ITS), and digital twins. Manual annotation of lanes is labor-intensive and costly, prompting researchers to explore automatic lane extraction methods. This paper presents an end-to-end large-scale lane mapping method that considers both lane geometry and semantics. This study represents lane markings as polylines with uniformly sampled points and associated semantics, allowing for adaptation to varying lane shapes. Additionally, we propose an end-to-end network to extract lane polylines from mobile laser scanning (MLS) data, enabling the inference of vectorized lane instances without complex post-processing. The network consists of three components: a feature encoder, a column proposal generator, and a lane information decoder. The feature encoder encodes textual and structural information of lane markings to enhance the method’s robustness to data imperfections, such as varying lane intensity, uneven point density, and occlusion-induced incomplete data. The column proposal generator generates regions of interest for the subsequent decoder. Leveraging the embedded multi-scale features from the feature encoder, the lane decoder effectively predicts lane polylines and their associated semantics without requiring step-by-step conditional inference. Comprehensive experiments conducted on three lane datasets have demonstrated the performance of the proposed method, even in the presence of incomplete data and complex lane topology. Furthermore, the datasets used in this work, including source ground points, generated bird’s eye view (BEV) images, and annotations, will be publicly available with the publication of the paper. ...

Simultaneous Local-Global Feature Learning for 3D Object Detection in Indoor Point Clouds

Journal article (2024) - Mingqiang Wei, Baian Chen, Liangliang Nan, Haoran Xie, Lipeng Gu, Dening Lu, Fu Lee Wang, Qing Li
The acquisition of both local and global features from irregular point clouds is crucial for 3D object detection (3DOD). Current mainstream 3D detectors neglect significant local features during pooling operations or disregard many global features of the overall scene context. This paper proposes new techniques for simultaneously learning local-global features of scene point clouds to enhance 3DOD. Specifically, we propose an efficient 3DOD network in indoor point clouds, named SimLOG, which utilizes simultaneous local-global feature learning. SimLOG has two main contributions: a Dynamic Points Interaction (DPI) module to recover local features lost during pooling, and a Global Context Aggregation(GCA) module to aggregate multi-scale features from various layers of the encoder to improve scene context awareness. Unlike traditional local-global feature learning methods, our DPI and GCA modules are integrated into a single feature learning module, making it easily detachable and able to be incorporated into existing 3DOD networks to enhance their performance. SimLOG demonstrates superior performance over twenty competitors in terms of detection accuracy and robustness on both the SUN RGB-D and ScanNet V2 datasets. Specifically, SimLOG boosts the baseline VoteNet by 8.1% of mAP@0.25 on ScanNet V2 and by 3.9% of mAP@0.25 on SUN RGB-D. ...

Polyhedron-based graph neural network for 3D building reconstruction from point clouds

Journal article (2024) - Zhaiyu Chen, Yilei Shi, Liangliang Nan, Zhitong Xiong, Xiao Xiang Zhu
We present PolyGNN, a polyhedron-based graph neural network for 3D building reconstruction from point clouds. PolyGNN learns to assemble primitives obtained by polyhedral decomposition via graph node classification, achieving a watertight and compact reconstruction. To effectively represent arbitrary-shaped polyhedra in the neural network, we propose a skeleton-based sampling strategy to generate polyhedron-wise queries. These queries are then incorporated with inter-polyhedron adjacency to enhance the classification. PolyGNN is end-to-end optimizable and is designed to accommodate variable-size input points, polyhedra, and queries with an index-driven batching technique. To address the abstraction gap between existing city-building models and the underlying instances, and provide a fair evaluation of the proposed method, we develop our method on a large-scale synthetic dataset with well-defined ground truths of polyhedral labels. We further conduct a transferability analysis across cities and on real-world point clouds. Both qualitative and quantitative results demonstrate the effectiveness of our method, particularly its efficiency for large-scale reconstructions. ...
Journal article (2024) - Jianwei Guo, Haobo Qin, Yinchang Zhou, Xin Chen, Liangliang Nan, Hui Huang
Digitalization of large-scale urban scenes (in particular buildings) has been a long-standing open problem, which attributes to the challenges in data acquisition, such as incomplete scene coverage, lack of semantics, low efficiency, and low reliability in path planning. In this paper, we address these challenges in urban building reconstruction from aerial images, and we propose an effective workflow and a few novel algorithms for efficient 3D building instance proxy reconstruction for large urban scenes. Specifically, we propose a novel learning-based approach to instance segmentation of urban buildings from aerial images followed by a voting-based algorithm to fuse the multi-view instance information to a sparse point cloud (reconstructed using a standard Structure from Motion pipeline). Our method enables effective instance segmentation of the building instances from the point cloud. We also introduce a layer-based surface reconstruction method dedicated to the 3D reconstruction of building proxies from extremely sparse point clouds. Extensive experiments on both synthetic and real-world aerial images of large urban scenes have demonstrated the effectiveness of our approach. The generated scene proxy models can already provide a promising 3D surface representation of the buildings in large urban scenes, and when applied to aerial path planning, the instance-enhanced building proxy models can significantly improve data completeness and accuracy, yielding highly detailed 3D building models. ...
Multi-sensor object detection is an active research topic in automated driving, but the robustness of such detection models against missing sensor input (modality missing), e.g., due to a sudden sensor failure, is a critical problem which remains under-studied. In this work, we propose UniBEV, an end-to-end multi-modal 3D object detection framework designed for robustness against missing modalities: UniBEV can operate on LiDAR plus camera input, but also on LiDAR-only or camera-only input without retraining. To facilitate its detector head to handle different input combinations, UniBEV aims to create well-aligned Bird’s Eye View (BEV) feature maps from each available modality. Unlike prior BEV-based multi-modal detection methods, all sensor modalities follow a uniform approach to resample features from the original sensor coordinate systems to the BEV features. We furthermore investigate the robustness of various fusion strategies w.r.t. missing modalities: the commonly used feature concatenation, but also channel-wise averaging, and a generalization to weighted averaging termed Channel Normalized Weights. To validate its effectiveness, we compare UniBEV to state-of-the-art BEVFusion and MetaBEV on nuScenes over all sensor input combinations. In this setting, UniBEV achieves better performance than these baselines for all input combinations. An ablation study shows the robustness benefits of fusing by weighted averaging over regular concatenation, and of sharing queries between the BEV encoders of each modality. Our code is available at https://github.com/tudelft-iv/UniBEV. ...