R.C. Lindenbergh
Please Note
77 records found
1
The Wadden in Photos
A morphological analysis from action camera pictures
This thesis focuses on a GoPro action camera installed on the mudflat near Holwerd. Over a three-month period, the camera took a picture every 15 minutes. The images show the changing mudflat, slowly transforming from a rolling landscape into a defined terrain. Analysis of the changing mudflat supports the research into morphology and sediment behaviour.
A Multilayer Perceptron (MLP) was constructed to perform semantic segmentation, where every pixel of the GoPro images is classified as either mud or water. A comprehensive model was set up, where manually labelled data and multiple image features were used to train the model and process the large image dataset. Model evaluation returned a macro-IoU of 0.71 and a macro-F1-score of 0.82. However, a detailed look at the classification results indicated critical model shortcomings and unnatural proportions of water and mud. Filtering of incorrect predictions resulted in a small dataset appropriate for further analysis.
The results indicate a strong correlation between predicted mud percentage and potential evaporation, revealing that this relationship can be observed from pictures. Spaghetti plots illustrated no patterns of change during single low-water periods. Finally, analysis of weekly average predictions reveals regions of growth and decrease on the mudflat. The channels and shallow pools are observed in particular, since these areas show divergent behaviour. It is theorised that the flow velocity influences the erosive and settling capacities of sediment. The shape of channels and shallow pools influences this velocity and, by extension, the morphological development of the mudflat.
Overall, the thesis demonstrates how an action camera on a mudflat can be used for observing both morphological changes and the forces that define this change. While the MLP does not deliver optimal results, it lays the foundation for future work on more advanced machine learning techniques. Finally, practical recommendations for future camera monitoring projects are given. ...
This thesis focuses on a GoPro action camera installed on the mudflat near Holwerd. Over a three-month period, the camera took a picture every 15 minutes. The images show the changing mudflat, slowly transforming from a rolling landscape into a defined terrain. Analysis of the changing mudflat supports the research into morphology and sediment behaviour.
A Multilayer Perceptron (MLP) was constructed to perform semantic segmentation, where every pixel of the GoPro images is classified as either mud or water. A comprehensive model was set up, where manually labelled data and multiple image features were used to train the model and process the large image dataset. Model evaluation returned a macro-IoU of 0.71 and a macro-F1-score of 0.82. However, a detailed look at the classification results indicated critical model shortcomings and unnatural proportions of water and mud. Filtering of incorrect predictions resulted in a small dataset appropriate for further analysis.
The results indicate a strong correlation between predicted mud percentage and potential evaporation, revealing that this relationship can be observed from pictures. Spaghetti plots illustrated no patterns of change during single low-water periods. Finally, analysis of weekly average predictions reveals regions of growth and decrease on the mudflat. The channels and shallow pools are observed in particular, since these areas show divergent behaviour. It is theorised that the flow velocity influences the erosive and settling capacities of sediment. The shape of channels and shallow pools influences this velocity and, by extension, the morphological development of the mudflat.
Overall, the thesis demonstrates how an action camera on a mudflat can be used for observing both morphological changes and the forces that define this change. While the MLP does not deliver optimal results, it lays the foundation for future work on more advanced machine learning techniques. Finally, practical recommendations for future camera monitoring projects are given.
Seabed Fingerprinting for Maritime Navigation in GNSS-Denied Environments
SAND-E: Seabed-Aided Navigation Using Classical and Learned Image Matching
This study investigates how diverse floodplain forest structure governs wave attenuation and evaluates Terrestrial Laser Scanning (TLS) as a method for deriving vegetation structural parameters across contrasting forest stands. The performance of TLS was evaluated using a reliability framework that defined the maximum distance over which vegetation structure could reliably be extracted from the point clouds.
Within this reliable domain, frontal surface area profiles (a(z)) were reconstructed and implemented in the phase-averaged wave model SWAN to simulate wave attenuation under varying water levels and wave forcing. Attenuation was governed by the interaction between submerged vegetation structure (a(z) Cd(z)) and wave orbital velocities (u(z)), resulting in dissipation proportional to a(z) Cd(z) u(z)³.
Pioneer and managed stands were characterised by concentrated low vegetation structure, limited horizontal patchiness, and structurally similar trees. Under moderate inundation conditions (1.6–4.0 m water depth above the forest floor), these stands produced the strongest wave attenuation with the smallest range of outcomes, with a median of approximately 40% and an interquartile range of 25–55% for a forest width of 100 m.
Late-successional stands, in contrast, were characterised by vertically distributed vegetation structure, pronounced horizontal patchiness, and structurally complex, diverse trees. Under the same conditions, these stands produced lower and more variable attenuation, with a median of approximately 20% and an interquartile range of 5–55%.
These results indicate that vertical vegetation structure primarily controls the magnitude of wave attenuation, whereas horizontal patchiness governs the variability of attenuation within forest stands. However, attenuation varied across hydraulic conditions and vegetation types, indicating that wave attenuation is not a fixed property of forest structure.
Compared to overly simplified, vertically uniform vegetation representations, TLS-derived a(z) profiles improved structural realism and captured depth-dependent attenuation behaviour. By linking high-resolution TLS-derived vegetation structure to wave modelling, this study provides a quantitative framework for evaluating structurally diverse floodplain forests as nature-based flood defences. ...
This study investigates how diverse floodplain forest structure governs wave attenuation and evaluates Terrestrial Laser Scanning (TLS) as a method for deriving vegetation structural parameters across contrasting forest stands. The performance of TLS was evaluated using a reliability framework that defined the maximum distance over which vegetation structure could reliably be extracted from the point clouds.
Within this reliable domain, frontal surface area profiles (a(z)) were reconstructed and implemented in the phase-averaged wave model SWAN to simulate wave attenuation under varying water levels and wave forcing. Attenuation was governed by the interaction between submerged vegetation structure (a(z) Cd(z)) and wave orbital velocities (u(z)), resulting in dissipation proportional to a(z) Cd(z) u(z)³.
Pioneer and managed stands were characterised by concentrated low vegetation structure, limited horizontal patchiness, and structurally similar trees. Under moderate inundation conditions (1.6–4.0 m water depth above the forest floor), these stands produced the strongest wave attenuation with the smallest range of outcomes, with a median of approximately 40% and an interquartile range of 25–55% for a forest width of 100 m.
Late-successional stands, in contrast, were characterised by vertically distributed vegetation structure, pronounced horizontal patchiness, and structurally complex, diverse trees. Under the same conditions, these stands produced lower and more variable attenuation, with a median of approximately 20% and an interquartile range of 5–55%.
These results indicate that vertical vegetation structure primarily controls the magnitude of wave attenuation, whereas horizontal patchiness governs the variability of attenuation within forest stands. However, attenuation varied across hydraulic conditions and vegetation types, indicating that wave attenuation is not a fixed property of forest structure.
Compared to overly simplified, vertically uniform vegetation representations, TLS-derived a(z) profiles improved structural realism and captured depth-dependent attenuation behaviour. By linking high-resolution TLS-derived vegetation structure to wave modelling, this study provides a quantitative framework for evaluating structurally diverse floodplain forests as nature-based flood defences.
Antarctic Time Machine
3D Reconstruction of Glaciers in the Antarctic Peninsula using Historical Structure-from-Motion
The workflow begins with semantic segmentation using a custom-trained U-Net model, which classifies pixels in degraded grayscale aerial images into six categories: snow, ice, water, rock, clouds, and sky. Despite limited training data (100 manually labelled images) and challenges such as low contrast and artefacts, the model achieves an overall accuracy of 73% and an F1-score of 71%. By masking out unusable regions such as sky and ocean, this step significantly improves the reliability of the photogrammetric reconstruction.
An automated metadata extraction module complements the segmentation by retrieving key parameters, including focal length, altitude, and fiducial marker positions, directly from the images. Using a combination of optical character recognition and computer vision techniques, it recovers essential information and estimates missing values by exploiting redundancy across flight series. This reduces the need for manual transcription and converts handwritten image annotations into structured digital formats.
The geo-referencing component establishes a spatial link between historical images and modern coordinate systems. It uses LightGlue, a recent deep-learning-based matching algorithm, along with a progressive tiling strategy adapted to the characteristics of historical imagery. By matching tie points between the TMA scans and Sentinel-2 satellite imagery, the system automatically generates ground control points (GCPs) with positional accuracies of just a few meters, therefore dramatically improving upon the original, often kilometre-scale geolocation estimates.
In the final stage, the segmented images, extracted metadata, and GCPs are automatically passed to Agisoft Metashape, which is integrated into the processing pipeline via its Python API. This stage performs Structurefrom- Motion photogrammetry to generate dense point clouds, orthophotos, and digital elevation models (DEMs) without user interaction. Applied across the Antarctic Peninsula, the pipeline successfully reconstructed 3D glacier surfaces for 49 glacier systems. Validation against the high-resolution Reference Elevation Model of Antarctica (REMA) shows median elevation differences of approximately 90 meters across full glacier extents and 76 meters in topographically stable areas.
While the outputs do not yet match the accuracy of fully manual processing, the developed system enables large-scale, repeatable reconstruction of historical glacier surfaces at a scale previously unattainable. By combining all components into a modular, end-to-end framework, this work makes the TMA archive broadly accessible for contemporary cryospheric research and extends observational baselines by over half a century. All code and workflows are openly available on GitHub, and the resulting data products, including semantic masks, metadata tables, geo-referenced image positions, and 3D glacier models, are publicly released to support further scientific use.
...
The workflow begins with semantic segmentation using a custom-trained U-Net model, which classifies pixels in degraded grayscale aerial images into six categories: snow, ice, water, rock, clouds, and sky. Despite limited training data (100 manually labelled images) and challenges such as low contrast and artefacts, the model achieves an overall accuracy of 73% and an F1-score of 71%. By masking out unusable regions such as sky and ocean, this step significantly improves the reliability of the photogrammetric reconstruction.
An automated metadata extraction module complements the segmentation by retrieving key parameters, including focal length, altitude, and fiducial marker positions, directly from the images. Using a combination of optical character recognition and computer vision techniques, it recovers essential information and estimates missing values by exploiting redundancy across flight series. This reduces the need for manual transcription and converts handwritten image annotations into structured digital formats.
The geo-referencing component establishes a spatial link between historical images and modern coordinate systems. It uses LightGlue, a recent deep-learning-based matching algorithm, along with a progressive tiling strategy adapted to the characteristics of historical imagery. By matching tie points between the TMA scans and Sentinel-2 satellite imagery, the system automatically generates ground control points (GCPs) with positional accuracies of just a few meters, therefore dramatically improving upon the original, often kilometre-scale geolocation estimates.
In the final stage, the segmented images, extracted metadata, and GCPs are automatically passed to Agisoft Metashape, which is integrated into the processing pipeline via its Python API. This stage performs Structurefrom- Motion photogrammetry to generate dense point clouds, orthophotos, and digital elevation models (DEMs) without user interaction. Applied across the Antarctic Peninsula, the pipeline successfully reconstructed 3D glacier surfaces for 49 glacier systems. Validation against the high-resolution Reference Elevation Model of Antarctica (REMA) shows median elevation differences of approximately 90 meters across full glacier extents and 76 meters in topographically stable areas.
While the outputs do not yet match the accuracy of fully manual processing, the developed system enables large-scale, repeatable reconstruction of historical glacier surfaces at a scale previously unattainable. By combining all components into a modular, end-to-end framework, this work makes the TMA archive broadly accessible for contemporary cryospheric research and extends observational baselines by over half a century. All code and workflows are openly available on GitHub, and the resulting data products, including semantic masks, metadata tables, geo-referenced image positions, and 3D glacier models, are publicly released to support further scientific use.
The reconstruction pipeline employed Structure from Motion and Multi-View Stereo algorithms, followed by a two stage point cloud registration. Coarse alignment was achieved through Sample Consensus Initial, and fine registration used the Iterative Closest Point algorithm. In the two benchmark datasets, registration achieved sub-metre mean C2C distances: 0.790 m for Toronto, 0.626 m for Westkapelle. In the test dataset, the mean C2C distance improved from 10.287 m before registration to 2.025 m afterwards, with 95% of points within 6 m. DEM differencing, supported by JARKUS cross-shore transect profiles, revealed systematic elevation gains of up to +10 m along foredune ridges, primarily resulting from a combination of documented coastal nourishment and natural process between 1990 and 2020.
However, limitations of the historical dataset, including sparse image coverage, strongly oblique viewing geometry, lack of vertical imagery, and poor GCP distribution, introduced geometric distortions and inconsistencies. The resulting orthophoto contained substantial voids and warped features, particularly in urban areas and low texture surfaces, underscoring the challenges of dense stereo matching under suboptimal imaging conditions.
Despite these constraints, this study demonstrates that meaningful reconstructions of past coastal environments are achievable when supported by careful preprocessing, robust registration, and multi source validation. The proposed workflow offers a transferable approach for extracting geomorphic insights from historical imagery in other coastal settings. ...
The reconstruction pipeline employed Structure from Motion and Multi-View Stereo algorithms, followed by a two stage point cloud registration. Coarse alignment was achieved through Sample Consensus Initial, and fine registration used the Iterative Closest Point algorithm. In the two benchmark datasets, registration achieved sub-metre mean C2C distances: 0.790 m for Toronto, 0.626 m for Westkapelle. In the test dataset, the mean C2C distance improved from 10.287 m before registration to 2.025 m afterwards, with 95% of points within 6 m. DEM differencing, supported by JARKUS cross-shore transect profiles, revealed systematic elevation gains of up to +10 m along foredune ridges, primarily resulting from a combination of documented coastal nourishment and natural process between 1990 and 2020.
However, limitations of the historical dataset, including sparse image coverage, strongly oblique viewing geometry, lack of vertical imagery, and poor GCP distribution, introduced geometric distortions and inconsistencies. The resulting orthophoto contained substantial voids and warped features, particularly in urban areas and low texture surfaces, underscoring the challenges of dense stereo matching under suboptimal imaging conditions.
Despite these constraints, this study demonstrates that meaningful reconstructions of past coastal environments are achievable when supported by careful preprocessing, robust registration, and multi source validation. The proposed workflow offers a transferable approach for extracting geomorphic insights from historical imagery in other coastal settings.
Classifying Bulldozers and Large Dynamic Objects on Sandy Beach Using LiDAR Point Clouds
An approach by multidimensional feature assessment
This study presents a robust framework for automatically classifying bulldozers and other large dynamic objects from terrestrial laser scanning (TLS) point clouds, achieving a test accuracy of 92.5% with a k-Nearest Neighbours (k-NN) classifier. This framework integrates both 3D point cloud descriptors with 2D projection-based features. Starting from raw TLS data, object clusters are extracted and described using geometric features. For each object, 3D features including linearity, planarity, and verticality are computed; 2D raster-based descriptors, including footprint spread and height variation, are computed from XY and XZ plane projections. These features are aggregated to build an object-level dataset. At the same time, global descriptors of the horizontal and vertical extents are computed to capture the dimension information. Together, these features are then standardised at the object level to train supervised classifiers capable of distinguishing four object classes: 'large bulldozer', 'other bulldozer', 'tractor-trailer', and 'other'.
The evaluation reveals that the instance-based k-NN model consistently outperformed a Support Vector Machine (SVM), which proves less robust to class imbalance and dataset shift due to its reliance on a fixed global decision boundary. Feature importance analysis confirms that a combination of 3D descriptors capturing structural complexity (e.g., eigenentropy, omnivariance) and 2D projection features quantifying vertical profiles is the most discriminative. The framework's real-world applicability is validated on an independently and automatically segmented dataset, where the k-NN model maintains a high overall accuracy of 90.9%. This validation also highlights the classification performance's sensitivity to segmentation quality, as incomplete object data from partial occlusion predictably decreases accuracy
In conclusion, this study establishes a practical and reliable feature-based methodology for monitoring anthropogenic activity in dynamic coastal zones. Future work should focus on enhancing this framework by expanding the training dataset to improve robustness, implementing adaptive binning for 2D feature extraction to better handle scale variance, and integrating more advanced segmentation algorithms to enable a fully automated monitoring pipeline. Such improvements will further solidify the method's utility for long-term environmental monitoring and data-driven coastal management. ...
This study presents a robust framework for automatically classifying bulldozers and other large dynamic objects from terrestrial laser scanning (TLS) point clouds, achieving a test accuracy of 92.5% with a k-Nearest Neighbours (k-NN) classifier. This framework integrates both 3D point cloud descriptors with 2D projection-based features. Starting from raw TLS data, object clusters are extracted and described using geometric features. For each object, 3D features including linearity, planarity, and verticality are computed; 2D raster-based descriptors, including footprint spread and height variation, are computed from XY and XZ plane projections. These features are aggregated to build an object-level dataset. At the same time, global descriptors of the horizontal and vertical extents are computed to capture the dimension information. Together, these features are then standardised at the object level to train supervised classifiers capable of distinguishing four object classes: 'large bulldozer', 'other bulldozer', 'tractor-trailer', and 'other'.
The evaluation reveals that the instance-based k-NN model consistently outperformed a Support Vector Machine (SVM), which proves less robust to class imbalance and dataset shift due to its reliance on a fixed global decision boundary. Feature importance analysis confirms that a combination of 3D descriptors capturing structural complexity (e.g., eigenentropy, omnivariance) and 2D projection features quantifying vertical profiles is the most discriminative. The framework's real-world applicability is validated on an independently and automatically segmented dataset, where the k-NN model maintains a high overall accuracy of 90.9%. This validation also highlights the classification performance's sensitivity to segmentation quality, as incomplete object data from partial occlusion predictably decreases accuracy
In conclusion, this study establishes a practical and reliable feature-based methodology for monitoring anthropogenic activity in dynamic coastal zones. Future work should focus on enhancing this framework by expanding the training dataset to improve robustness, implementing adaptive binning for 2D feature extraction to better handle scale variance, and integrating more advanced segmentation algorithms to enable a fully automated monitoring pipeline. Such improvements will further solidify the method's utility for long-term environmental monitoring and data-driven coastal management.
Exploring the Valkenburg mines in virtual reality
Obtaining and processing 3D LiDAR data from the Valkenburg mines for use in virtual reality
Identifying Dynamic Objects in Coastal Environments
Automatic Detection and Clustering of Dynamic Objects in Sequential LiDAR Point Cloud Data of the Beach of Noordwijk
within a defined range and validated against manually identified dynamic objects. For a week-long dataset, the error in the number of detected large dynamic objects is relatively low at 6.9%, whereas the error for small dynamic objects is higher at 23.0%, attributed to their proximity to the ground and to each other. On a point-to-point basis, the optimized configuration results in an average error of 13.5% for large dynamic objects and 33.7% for small dynamic objects with respect to a reference set. A sensitivity analysis using a Monte Carlo simulation with normally distributed parameter variations around the tuned values demonstrates robustness to moderate parameter fluctuations, particularly for larger dynamic objects, which show a standard deviation of 0.07 detected objects, while smaller objects show greater variability with a standard deviation of 0.34 detected objects. The application of the Cloth Simulation Filter adds value by excluding geomorphological processes, contributing to a reduction in error rate of approximately 95% for large objects and 62% for small objects. Overall, the presented workflow offers a robust automated approach for detecting dynamic objects on sandy beaches in LiDAR point cloud data, with demonstrated potential for scalable, long-term monitoring. ...
within a defined range and validated against manually identified dynamic objects. For a week-long dataset, the error in the number of detected large dynamic objects is relatively low at 6.9%, whereas the error for small dynamic objects is higher at 23.0%, attributed to their proximity to the ground and to each other. On a point-to-point basis, the optimized configuration results in an average error of 13.5% for large dynamic objects and 33.7% for small dynamic objects with respect to a reference set. A sensitivity analysis using a Monte Carlo simulation with normally distributed parameter variations around the tuned values demonstrates robustness to moderate parameter fluctuations, particularly for larger dynamic objects, which show a standard deviation of 0.07 detected objects, while smaller objects show greater variability with a standard deviation of 0.34 detected objects. The application of the Cloth Simulation Filter adds value by excluding geomorphological processes, contributing to a reduction in error rate of approximately 95% for large objects and 62% for small objects. Overall, the presented workflow offers a robust automated approach for detecting dynamic objects on sandy beaches in LiDAR point cloud data, with demonstrated potential for scalable, long-term monitoring.
Machine learning techniques using convolutional neural network (CNN) models can be applied to the field of archaeology to search for excavation sites automatically. An obstacle remains: CNN models require a lot of training image data to be efficient at their task, which is to classify whether areas contain rondels or otherwise. There are only 20 visible rondels on the orthophotomosaic that can be used as training images for model input, creating an imbalanced data set of a class with a minority class of aerial images of rondels and a large majority class of aerial images without rondels. Sketches and recorded characteristics of rondels from current images and from archaeological publications were used to automatically and randomly replicate rondel appearances from above, resulting in a created balanced data set with sufficient rondel examples for CNN training.
Multiple ResNet-34 models and a ConvNeXt model with differing hyperparameters were trained. The most promising model, a modified ResNet-34, was selected based on validation loss from the cross- entropy loss function and on the number of correctly and incorrectly classified labeled images from a test set. The selected model is used to classify data from the orthophotomosaic for rondels using a sliding window technique. Over 9510 square kilometers of agricultural land cover in western and eastern Slovakia was selected from the CORINE land cover map for classification. 7 suspected rondel sites were found, and 2 were determined to likely be rondels, based on their circular ditch-like appearance in 4 sets of multispectral images and in LiDAR elevation data. Results indicate that exact rondel layouts can be delineated with high-resolution orthophotomasics, however identifying circular elevation patterns of ditches proves to be challenging without using additional LiDAR data. ...
Machine learning techniques using convolutional neural network (CNN) models can be applied to the field of archaeology to search for excavation sites automatically. An obstacle remains: CNN models require a lot of training image data to be efficient at their task, which is to classify whether areas contain rondels or otherwise. There are only 20 visible rondels on the orthophotomosaic that can be used as training images for model input, creating an imbalanced data set of a class with a minority class of aerial images of rondels and a large majority class of aerial images without rondels. Sketches and recorded characteristics of rondels from current images and from archaeological publications were used to automatically and randomly replicate rondel appearances from above, resulting in a created balanced data set with sufficient rondel examples for CNN training.
Multiple ResNet-34 models and a ConvNeXt model with differing hyperparameters were trained. The most promising model, a modified ResNet-34, was selected based on validation loss from the cross- entropy loss function and on the number of correctly and incorrectly classified labeled images from a test set. The selected model is used to classify data from the orthophotomosaic for rondels using a sliding window technique. Over 9510 square kilometers of agricultural land cover in western and eastern Slovakia was selected from the CORINE land cover map for classification. 7 suspected rondel sites were found, and 2 were determined to likely be rondels, based on their circular ditch-like appearance in 4 sets of multispectral images and in LiDAR elevation data. Results indicate that exact rondel layouts can be delineated with high-resolution orthophotomasics, however identifying circular elevation patterns of ditches proves to be challenging without using additional LiDAR data.
A series of controlled field-based experiments is conducted to measure field-of-view (FOV) coverage, point-density distribution, range precision, and sensitivity to vibrations. Python tools are used to estimate FOV coverage and density over time, while PCA-based plane fitting determines the distance random error at 20m. Additional analysis tests the influence of external forces causing vibrations and assess long-term stability using data from the Internal Measurement Unit (IMU). To achieve practical deployment, a portable Central Observations Recorder (COR) integrating a Raspberry Pi controller, power management, and anemometer connectivity is designed, developed and tested. Time series derived from point clouds are processed into 3D motion fields using PlantMove to analyse tree displacement patterns in order to showcase the AVIA's dynamic scanning capabilities.
Results show that the AVIA achieves approximately 92% FOV coverage within 1000 ms (contrasting the manufacturer’s 800 ms claim) and maintains high range precision (σ ≈ 0.8 cm at 20 m), exceeding stated specifications. Point density is found to be strongly non-uniform over the FOV, with the central half of the FOV exhibiting a roughly 2.4 times higher density. Small, irregular vibrations increase range noise by less than 1 mm, while airborne particles and heavy precipitation further reduce return intensity and point density. Scans made with hours of time in between showed negligible drift between them, confirming the sensor’s stability for longer-term monitoring setups.
Overall, the findings demonstrate that the Livox AVIA is a reliable, precise, and low-cost LiDAR sensor that can be used for near to mid range (2 – 100 m) static environmental monitoring. When appropriately configured and keeping in mind its non-uniform FOV density and full FOV coverage time, the system performs effectively in both manual and autonomous operations for monitoring dynamic processes. ...
A series of controlled field-based experiments is conducted to measure field-of-view (FOV) coverage, point-density distribution, range precision, and sensitivity to vibrations. Python tools are used to estimate FOV coverage and density over time, while PCA-based plane fitting determines the distance random error at 20m. Additional analysis tests the influence of external forces causing vibrations and assess long-term stability using data from the Internal Measurement Unit (IMU). To achieve practical deployment, a portable Central Observations Recorder (COR) integrating a Raspberry Pi controller, power management, and anemometer connectivity is designed, developed and tested. Time series derived from point clouds are processed into 3D motion fields using PlantMove to analyse tree displacement patterns in order to showcase the AVIA's dynamic scanning capabilities.
Results show that the AVIA achieves approximately 92% FOV coverage within 1000 ms (contrasting the manufacturer’s 800 ms claim) and maintains high range precision (σ ≈ 0.8 cm at 20 m), exceeding stated specifications. Point density is found to be strongly non-uniform over the FOV, with the central half of the FOV exhibiting a roughly 2.4 times higher density. Small, irregular vibrations increase range noise by less than 1 mm, while airborne particles and heavy precipitation further reduce return intensity and point density. Scans made with hours of time in between showed negligible drift between them, confirming the sensor’s stability for longer-term monitoring setups.
Overall, the findings demonstrate that the Livox AVIA is a reliable, precise, and low-cost LiDAR sensor that can be used for near to mid range (2 – 100 m) static environmental monitoring. When appropriately configured and keeping in mind its non-uniform FOV density and full FOV coverage time, the system performs effectively in both manual and autonomous operations for monitoring dynamic processes.
Urban Tree Classification in Delft, the Netherlands
Classifying Urban Tree Characteristics with Machine Learning Using Airborne LiDAR and Satellite Imagery
This study proposes a workflow for automating quality validation of LiDAR infrastructure point clouds. The workflow assesses point cloud quality based on three primary components: coverage, relative accuracy, and absolute accuracy. The methodology includes:
• Point Cloud Density assessment: involves analyzing 2D horizontal cells and partial-3D spaces to ensure compliance with density requirements.
• Overlapping Regions Alignment: identifies and compares surfaces in overlapping areas of point clouds to determine relative accuracy.
• Benchmark alignment: extracts points corresponding to spherical targets, estimates the center coordinates, and evaluates adherence to absolute accuracy standards.
Quality assessments were conducted on static, mobile, and airborne point clouds. The static point cloud analysis revealed non-compliance with density requirements in 2D, with approximately half the points failing to meet standards. In 3D analysis, compliance was observed for 1 m2 horizontal cells, but individual 1-meter sections often fell short upon closer inspection. These findings highlight the need for tailored quality standards: detailed 3D analysis is crucial for complex environments like tunnels, while road environments can be effectively evaluated in 2D. Relative accuracy assessments for static and mobile datasets showed compliance with scanner specifications, with RMSE values meeting the specified requirements. Absolute accuracy assessments on static point clouds met requirements with minimal deviations in both XY and Z directions.
Recommendations include defining density requirements for different environments, establishing acceptance criteria, and defining allowable deviations for all requirements. ...
This study proposes a workflow for automating quality validation of LiDAR infrastructure point clouds. The workflow assesses point cloud quality based on three primary components: coverage, relative accuracy, and absolute accuracy. The methodology includes:
• Point Cloud Density assessment: involves analyzing 2D horizontal cells and partial-3D spaces to ensure compliance with density requirements.
• Overlapping Regions Alignment: identifies and compares surfaces in overlapping areas of point clouds to determine relative accuracy.
• Benchmark alignment: extracts points corresponding to spherical targets, estimates the center coordinates, and evaluates adherence to absolute accuracy standards.
Quality assessments were conducted on static, mobile, and airborne point clouds. The static point cloud analysis revealed non-compliance with density requirements in 2D, with approximately half the points failing to meet standards. In 3D analysis, compliance was observed for 1 m2 horizontal cells, but individual 1-meter sections often fell short upon closer inspection. These findings highlight the need for tailored quality standards: detailed 3D analysis is crucial for complex environments like tunnels, while road environments can be effectively evaluated in 2D. Relative accuracy assessments for static and mobile datasets showed compliance with scanner specifications, with RMSE values meeting the specified requirements. Absolute accuracy assessments on static point clouds met requirements with minimal deviations in both XY and Z directions.
Recommendations include defining density requirements for different environments, establishing acceptance criteria, and defining allowable deviations for all requirements.
JARKUS data is obtained from airborne laser scanning for the Dutch coastal areas which is done yearly between January and March. Eight datasets of annual lidar data which include the dunes are used for this study. The datasets consists of point clouds for the years 2016 to 2023.
To obtain information about the changes occurring, several methods are used. The workflow recommended is the C2M method for 3D changes. These changes can be clustered by K-means clustering. 2D changes can be observed using cross sections, contour lines and the volume changes.
The seaward side of the dunes decreases in height due to marine erosion such as storms. Around 4.5 meter erosion on some locations has been found due to large storms in 2022. The lagoon side of the dunes increases in height by aeolian transport. The rate of dune growth is found to be 0.5 meters per year.
The middle dune experiences more erosion than deposition and the volume decreases. The erosion is caused by less vegetation and human impacts. If the trend from before the storm in 2022 continues, the dune will shrink and disappear if no additional maintenance is done. The other two dunes do not experience big volume losses and are classified as stable. ...
JARKUS data is obtained from airborne laser scanning for the Dutch coastal areas which is done yearly between January and March. Eight datasets of annual lidar data which include the dunes are used for this study. The datasets consists of point clouds for the years 2016 to 2023.
To obtain information about the changes occurring, several methods are used. The workflow recommended is the C2M method for 3D changes. These changes can be clustered by K-means clustering. 2D changes can be observed using cross sections, contour lines and the volume changes.
The seaward side of the dunes decreases in height due to marine erosion such as storms. Around 4.5 meter erosion on some locations has been found due to large storms in 2022. The lagoon side of the dunes increases in height by aeolian transport. The rate of dune growth is found to be 0.5 meters per year.
The middle dune experiences more erosion than deposition and the volume decreases. The erosion is caused by less vegetation and human impacts. If the trend from before the storm in 2022 continues, the dune will shrink and disappear if no additional maintenance is done. The other two dunes do not experience big volume losses and are classified as stable.
The methodology emphasizes the modification of existing pipelines to optimize the handling of satellite images, incorporating sophisticated feature detection and matching technologies. The performance of these adaptations is rigorously evaluated through extensive analysis using satellite imagery across varied resolutions and environmental conditions. Results from the study indicate marked improvements in the fidelity and accuracy of the generated DEMs, which are substantiated by validation against high-resolution LiDAR ground truth data.
The refined pipeline effectively manages multi-source satellite images and produces terrain models of significantly higher quality, vital for robust geospatial analysis. This work not only bridges the gap between remote sensing and computer vision but also lays the groundwork for future research aimed at improving DEM generation from satellite imagery. This study proposes potential transformative practices in geospatial analysis and supports continued progress in fields such as environmental monitoring, urban planning, and disaster management. ...
The methodology emphasizes the modification of existing pipelines to optimize the handling of satellite images, incorporating sophisticated feature detection and matching technologies. The performance of these adaptations is rigorously evaluated through extensive analysis using satellite imagery across varied resolutions and environmental conditions. Results from the study indicate marked improvements in the fidelity and accuracy of the generated DEMs, which are substantiated by validation against high-resolution LiDAR ground truth data.
The refined pipeline effectively manages multi-source satellite images and produces terrain models of significantly higher quality, vital for robust geospatial analysis. This work not only bridges the gap between remote sensing and computer vision but also lays the groundwork for future research aimed at improving DEM generation from satellite imagery. This study proposes potential transformative practices in geospatial analysis and supports continued progress in fields such as environmental monitoring, urban planning, and disaster management.
Deep Learning-based Segmentation of Cracks within a Photogrammetry Solution
Fully-Supervised Learning, Transfer Learning and Photogrammetric Image Processing
As manual visual inspection of this imagery is very time-consuming, this work proposes a methodology based on fully-supervised deep learning-based segmentation techniques with the goal of detecting and localizing cracks in the masonry quay walls. For this purpose, two neural networks are trained, one for the segmentation of quay walls in images, and one for the segmentation of cracks.
The neural network architectures which are considered in this work are DeepLabV3+, FPN, MANet and LinkNet, together with different encoders and loss functions. For quay wall segmentation, we adopt transfer learning on a network trained on masonry walls and fine-tune it for quay walls specifically. Here, DeepLabV3+ with ResNeXt-50 was found to be most effective, achieving a F1-score of 96.3 % on the test set. For crack segmentation, FPN with ResNeSt-50 performed best, resulting in a test set F1-score of 78.8 %.
The inference of the crack network is done with a multi-level scheme to detect cracks at different image scales and increase output confidence.
The inherent photogrammetric properties of the imagery have proven to be vital for further post-processing steps, like aggregating overlapping predictions, resulting in more prediction confidence.
Photogrammetry also enables converting pixel-wise predictions to crack length and crack width in the units of meters and millimeters respectively. The methodology additionally proposes photogrammetric image processing methods to transform neural network predictions to a 3D representation and a true-to-scale orthographic 2D image.
Additionally a concise visual evaluation has been conducted to assess the prediction performance on an otherwise unlabelled dataset.
This thesis presents an engineering effort for fully-supervised crack localization within the context of photogrammetric processed images, with generalization in mind for automatic assessment. ...
As manual visual inspection of this imagery is very time-consuming, this work proposes a methodology based on fully-supervised deep learning-based segmentation techniques with the goal of detecting and localizing cracks in the masonry quay walls. For this purpose, two neural networks are trained, one for the segmentation of quay walls in images, and one for the segmentation of cracks.
The neural network architectures which are considered in this work are DeepLabV3+, FPN, MANet and LinkNet, together with different encoders and loss functions. For quay wall segmentation, we adopt transfer learning on a network trained on masonry walls and fine-tune it for quay walls specifically. Here, DeepLabV3+ with ResNeXt-50 was found to be most effective, achieving a F1-score of 96.3 % on the test set. For crack segmentation, FPN with ResNeSt-50 performed best, resulting in a test set F1-score of 78.8 %.
The inference of the crack network is done with a multi-level scheme to detect cracks at different image scales and increase output confidence.
The inherent photogrammetric properties of the imagery have proven to be vital for further post-processing steps, like aggregating overlapping predictions, resulting in more prediction confidence.
Photogrammetry also enables converting pixel-wise predictions to crack length and crack width in the units of meters and millimeters respectively. The methodology additionally proposes photogrammetric image processing methods to transform neural network predictions to a 3D representation and a true-to-scale orthographic 2D image.
Additionally a concise visual evaluation has been conducted to assess the prediction performance on an otherwise unlabelled dataset.
This thesis presents an engineering effort for fully-supervised crack localization within the context of photogrammetric processed images, with generalization in mind for automatic assessment.
Thus, the aim of this research is to develop an improved trajectory optimization method, thereby ensuring accurate geo-referencing and alignment of the survey data. This thesis proposes a newly developed methodology to achieve this aim: features are extracted from point cloud surveys, matched and utilized by g2o optimizer and GNSS processing software to optimize the trajectory. The development is described and results are evaluated on two different scales - locally, within a point cloud tile and globally, within a sequence of tiles. It is done by using Glasgow's underground railway network as a test case.
Results from the implementation demonstrate significant improvements in trajectory accuracy - a misalignment of point cloud data was reduced from a 1.5 m to a cm level within an optimization time frame that took approximately 10 hours. This improvement in accuracy was present under different complex environments using both the local and global versions of the algorithm. However, the area near the railway tunnel entrance saw a limited benefit from the implementation of the proposed algorithm.
In conclusion, the developed trajectory optimization algorithm optimizes the trajectory and improves the alignment of the survey data. Moreover, the method outperforms the currently employed solutions by being automatic and applicable in different environments. However, further research is required to optimize the algorithm itself (accuracy and computationally speed of the algorithm) and to more accurately define its limitations in terms of the surveyed environments. ...
Thus, the aim of this research is to develop an improved trajectory optimization method, thereby ensuring accurate geo-referencing and alignment of the survey data. This thesis proposes a newly developed methodology to achieve this aim: features are extracted from point cloud surveys, matched and utilized by g2o optimizer and GNSS processing software to optimize the trajectory. The development is described and results are evaluated on two different scales - locally, within a point cloud tile and globally, within a sequence of tiles. It is done by using Glasgow's underground railway network as a test case.
Results from the implementation demonstrate significant improvements in trajectory accuracy - a misalignment of point cloud data was reduced from a 1.5 m to a cm level within an optimization time frame that took approximately 10 hours. This improvement in accuracy was present under different complex environments using both the local and global versions of the algorithm. However, the area near the railway tunnel entrance saw a limited benefit from the implementation of the proposed algorithm.
In conclusion, the developed trajectory optimization algorithm optimizes the trajectory and improves the alignment of the survey data. Moreover, the method outperforms the currently employed solutions by being automatic and applicable in different environments. However, further research is required to optimize the algorithm itself (accuracy and computationally speed of the algorithm) and to more accurately define its limitations in terms of the surveyed environments.
Geo-localization is achieved by comparing the historical image with positions within a predefined geo-referenced Area of Interest (AoI). Two predefined remote sensing datasets are used: Sentinel-2 and Quantarctica Rock Outcrop Mask, from which AoIs are generated. Positions within the AoI exhibiting the highest similarity to the historical image are likely to correspond to the same ground area, thus providing the location of the historical imagery.
This similarity assessment employs two Siamese Networks: SigNet and ResNet-50. SigNet, originally designed for signature verification tasks, consists of four convolutional layers. In contrast, ResNet-50, initially developed for image classification purposes, is characterized by its deep architecture comprising approximately 50 convolutional layers, as suggested by its name. In this study, these two models are initially pre-trained on cross-domain datasets and subsequently adaptively trained with task-specific datasets created in this study. The adaptive training datasets comprise triplets of similar and dissimilar images pre-processed using methods devised in this study. An evaluation methodology based on confidence level is developed to assess the model and workflow performance, which is then applied to 51 test historical image samples.
Overall, the results indicate that the ResNet-50 based network outperforms SigNet, achieving a 95.5% average confidence level. However, the method does not meet the initial expectation of directly providing the location of the historical image within the AoI. Instead, it identifies potential locations. Nevertheless, this outcome is valuable as it streamlines the search process for subsequent image matching steps. For instance, a 95.5% average confidence level for the ResNet-50 based network correlates with an approximate 95.5% reduction in processing time for geo-referencing when integrated with image matching in subsequent steps. ...
Geo-localization is achieved by comparing the historical image with positions within a predefined geo-referenced Area of Interest (AoI). Two predefined remote sensing datasets are used: Sentinel-2 and Quantarctica Rock Outcrop Mask, from which AoIs are generated. Positions within the AoI exhibiting the highest similarity to the historical image are likely to correspond to the same ground area, thus providing the location of the historical imagery.
This similarity assessment employs two Siamese Networks: SigNet and ResNet-50. SigNet, originally designed for signature verification tasks, consists of four convolutional layers. In contrast, ResNet-50, initially developed for image classification purposes, is characterized by its deep architecture comprising approximately 50 convolutional layers, as suggested by its name. In this study, these two models are initially pre-trained on cross-domain datasets and subsequently adaptively trained with task-specific datasets created in this study. The adaptive training datasets comprise triplets of similar and dissimilar images pre-processed using methods devised in this study. An evaluation methodology based on confidence level is developed to assess the model and workflow performance, which is then applied to 51 test historical image samples.
Overall, the results indicate that the ResNet-50 based network outperforms SigNet, achieving a 95.5% average confidence level. However, the method does not meet the initial expectation of directly providing the location of the historical image within the AoI. Instead, it identifies potential locations. Nevertheless, this outcome is valuable as it streamlines the search process for subsequent image matching steps. For instance, a 95.5% average confidence level for the ResNet-50 based network correlates with an approximate 95.5% reduction in processing time for geo-referencing when integrated with image matching in subsequent steps.