EP
E.A.I. Pool
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
4 records found
1
This master thesis presents an experimental study on 3D person localization (i.e., pedestrians, cyclists)in traffic scenes, using monocular vision and Light Detection And Ranging (LiDAR) data. The performance of two top-ranking methods is analyzed on the 3D object detection KITTI dataset. In this evaluation, the effect of the Intersection over Union (IoU) threshold on the performance in terms of 3D bounding box location, size, and orientation is analysed.
Since the KITTI 3D object detection dataset contains relatively few 3D person instances, the analysis will is to the EuroCity Persons 2.5D (ECP2.5D) datasets (both day and night), which is one order of magnitude larger. Using both datasets, additional experiments are performed to evaluate the influence of distance, the number of LiDAR points, occlusion, and intensity on the performance. Domain transfer experiments between the KITTI and ECP2.5D datasets are performed, to examine how these datasets generalize with respect to each other. Furthermore, Part-A2 net is used to evaluate the detection score which is given to the ground truth pedestrians. The relationship between the detection score and the distance, the number of LiDAR points, and occlusion is analyzed. Some objects are not detected although their ground truth detection score is high. This creates the potential to detect these pedestrians. Lastly, this thesis presents a method that uses the detections from the previous frame to increase the performance in the subsequent frame by adding the previous detections to the 3D proposals coming from the Region Proposal Network (RPN). ...
Since the KITTI 3D object detection dataset contains relatively few 3D person instances, the analysis will is to the EuroCity Persons 2.5D (ECP2.5D) datasets (both day and night), which is one order of magnitude larger. Using both datasets, additional experiments are performed to evaluate the influence of distance, the number of LiDAR points, occlusion, and intensity on the performance. Domain transfer experiments between the KITTI and ECP2.5D datasets are performed, to examine how these datasets generalize with respect to each other. Furthermore, Part-A2 net is used to evaluate the detection score which is given to the ground truth pedestrians. The relationship between the detection score and the distance, the number of LiDAR points, and occlusion is analyzed. Some objects are not detected although their ground truth detection score is high. This creates the potential to detect these pedestrians. Lastly, this thesis presents a method that uses the detections from the previous frame to increase the performance in the subsequent frame by adding the previous detections to the 3D proposals coming from the Region Proposal Network (RPN). ...
This master thesis presents an experimental study on 3D person localization (i.e., pedestrians, cyclists)in traffic scenes, using monocular vision and Light Detection And Ranging (LiDAR) data. The performance of two top-ranking methods is analyzed on the 3D object detection KITTI dataset. In this evaluation, the effect of the Intersection over Union (IoU) threshold on the performance in terms of 3D bounding box location, size, and orientation is analysed.
Since the KITTI 3D object detection dataset contains relatively few 3D person instances, the analysis will is to the EuroCity Persons 2.5D (ECP2.5D) datasets (both day and night), which is one order of magnitude larger. Using both datasets, additional experiments are performed to evaluate the influence of distance, the number of LiDAR points, occlusion, and intensity on the performance. Domain transfer experiments between the KITTI and ECP2.5D datasets are performed, to examine how these datasets generalize with respect to each other. Furthermore, Part-A2 net is used to evaluate the detection score which is given to the ground truth pedestrians. The relationship between the detection score and the distance, the number of LiDAR points, and occlusion is analyzed. Some objects are not detected although their ground truth detection score is high. This creates the potential to detect these pedestrians. Lastly, this thesis presents a method that uses the detections from the previous frame to increase the performance in the subsequent frame by adding the previous detections to the 3D proposals coming from the Region Proposal Network (RPN).
Since the KITTI 3D object detection dataset contains relatively few 3D person instances, the analysis will is to the EuroCity Persons 2.5D (ECP2.5D) datasets (both day and night), which is one order of magnitude larger. Using both datasets, additional experiments are performed to evaluate the influence of distance, the number of LiDAR points, occlusion, and intensity on the performance. Domain transfer experiments between the KITTI and ECP2.5D datasets are performed, to examine how these datasets generalize with respect to each other. Furthermore, Part-A2 net is used to evaluate the detection score which is given to the ground truth pedestrians. The relationship between the detection score and the distance, the number of LiDAR points, and occlusion is analyzed. Some objects are not detected although their ground truth detection score is high. This creates the potential to detect these pedestrians. Lastly, this thesis presents a method that uses the detections from the previous frame to increase the performance in the subsequent frame by adding the previous detections to the 3D proposals coming from the Region Proposal Network (RPN).
A human driver can gauge the intention and signals given by other road users indicative of their future behaviour. The intentions and signals are identified by looking at the cues originating from vulnerable road users or their surroundings (hand signals, head orientation, posture, traffic signals, distance to curb, etc.). Taking all these cues into account by creating a separate detector for each is an extremely difficult task. Instead, this MSc Thesis will explore the possibility of using a generic contextual cue in optical flow originating from a pedestrian with deep learning methods to improve the path prediction in a naturalistic driving scenario. The contribution of this work is to examine multiple ways to extract relevant information from the optical flow and also explore the possibility of using the entire the high-dimensional optical flow using convolutions and soft-attention to help identify relevant pixels for the prediction task. This work elaborates on the extraction and processing of optical flow features. It proposes 2 Recurrent neural networks (RNN) based model: one to work with the histogram of optical flow features and the other one to take in the dense optical flow directly. Also, visualization of the soft-attention weights is done to add a step that helps in the interpretability of the RNN model incorporating dense optical flow. From the experimental results, optical flow features have shown significant improvements in terms of predicting probabilistic confidence for tracks with some changes in their motion mode. It was seen that the convolution-attention RNN model was able to work with dense optical flow features and position of pedestrians as input to obtain better results among all the combinations of features and models compared in this work.
...
A human driver can gauge the intention and signals given by other road users indicative of their future behaviour. The intentions and signals are identified by looking at the cues originating from vulnerable road users or their surroundings (hand signals, head orientation, posture, traffic signals, distance to curb, etc.). Taking all these cues into account by creating a separate detector for each is an extremely difficult task. Instead, this MSc Thesis will explore the possibility of using a generic contextual cue in optical flow originating from a pedestrian with deep learning methods to improve the path prediction in a naturalistic driving scenario. The contribution of this work is to examine multiple ways to extract relevant information from the optical flow and also explore the possibility of using the entire the high-dimensional optical flow using convolutions and soft-attention to help identify relevant pixels for the prediction task. This work elaborates on the extraction and processing of optical flow features. It proposes 2 Recurrent neural networks (RNN) based model: one to work with the histogram of optical flow features and the other one to take in the dense optical flow directly. Also, visualization of the soft-attention weights is done to add a step that helps in the interpretability of the RNN model incorporating dense optical flow. From the experimental results, optical flow features have shown significant improvements in terms of predicting probabilistic confidence for tracks with some changes in their motion mode. It was seen that the convolution-attention RNN model was able to work with dense optical flow features and position of pedestrians as input to obtain better results among all the combinations of features and models compared in this work.
Classification of Damages on Aircraft Inspection Images Using Convolutional Neural Networks
Kick-starting a Deep Learning project with limited data
Master thesis
(2019)
-
Julian Freiherr von der Goltz, Wei Pan, Jens Kober, Ewoud Pool, Alfredo Nunez, Jochem Verboom
Aircraft inspections after unexpected incidents, like lightning strikes, currently require a timeconsuming and costly inspection process, due to the small size of the lightning strike damages. Mainblades Inspections is working on an automated, drone-based solution, that scans the aircraft hull with a high-resolution camera. The objective of this project is to assess the
feasibility of using a deep Convolutional Neural Network (CNN) for (semi-)automated damage detection, with the goal of achieving a high recall (low False Negative Rate (FNR)) on the small damages. The problem is framed as a classification problem on limited and imbalanced data. However, it is not pre-defined if single-label or multi-label classification should be used,
and both approaches are investigated.
The main contribution of this work is to show experimentally how common deep CNN architectures and Deep Learning practices can be used to train classifiers that recognize damages in a specialized domain, with application-specific metrics. We present methods for synthesizing, pre-processing and re-sampling of the necessary dataset. It is shown that pre-trained, parameter-efficient CNN architectures that implement skip-connections, complemented by
global max-pooling before the final layer, are well suited for that dataset. The Xception architecture has been chosen as backbone for the classifier due to its high recall and fast convergence. To mitigate the detrimental influence of imbalanced training data, training data re-sampling that equalizes the class distribution is implemented. It has a positive effect recall, especially
when applied to multi-label classification. When using re-sampling and data augmentation, the performance of multi-label and single-label classification can be brought to the same level. However, the best achieved FNR is 5.4%, with a softmax classifier, combining all regularization methods.
Finally, we investigated how regularization can be used to increase generalization capability with limited training data. Data augmentation is the most effective regularization method, even though its full potential has not been explored yet. Dropout benefits single-label classification
but not multi-label classification. L2-regularization has a moderate positive effect on both. Naively combining the regularization techniques without an exhaustive grid search or automated search on average does not yield any additional gains and shows the limit of manual hyper-parameter tuning. ...
feasibility of using a deep Convolutional Neural Network (CNN) for (semi-)automated damage detection, with the goal of achieving a high recall (low False Negative Rate (FNR)) on the small damages. The problem is framed as a classification problem on limited and imbalanced data. However, it is not pre-defined if single-label or multi-label classification should be used,
and both approaches are investigated.
The main contribution of this work is to show experimentally how common deep CNN architectures and Deep Learning practices can be used to train classifiers that recognize damages in a specialized domain, with application-specific metrics. We present methods for synthesizing, pre-processing and re-sampling of the necessary dataset. It is shown that pre-trained, parameter-efficient CNN architectures that implement skip-connections, complemented by
global max-pooling before the final layer, are well suited for that dataset. The Xception architecture has been chosen as backbone for the classifier due to its high recall and fast convergence. To mitigate the detrimental influence of imbalanced training data, training data re-sampling that equalizes the class distribution is implemented. It has a positive effect recall, especially
when applied to multi-label classification. When using re-sampling and data augmentation, the performance of multi-label and single-label classification can be brought to the same level. However, the best achieved FNR is 5.4%, with a softmax classifier, combining all regularization methods.
Finally, we investigated how regularization can be used to increase generalization capability with limited training data. Data augmentation is the most effective regularization method, even though its full potential has not been explored yet. Dropout benefits single-label classification
but not multi-label classification. L2-regularization has a moderate positive effect on both. Naively combining the regularization techniques without an exhaustive grid search or automated search on average does not yield any additional gains and shows the limit of manual hyper-parameter tuning. ...
Aircraft inspections after unexpected incidents, like lightning strikes, currently require a timeconsuming and costly inspection process, due to the small size of the lightning strike damages. Mainblades Inspections is working on an automated, drone-based solution, that scans the aircraft hull with a high-resolution camera. The objective of this project is to assess the
feasibility of using a deep Convolutional Neural Network (CNN) for (semi-)automated damage detection, with the goal of achieving a high recall (low False Negative Rate (FNR)) on the small damages. The problem is framed as a classification problem on limited and imbalanced data. However, it is not pre-defined if single-label or multi-label classification should be used,
and both approaches are investigated.
The main contribution of this work is to show experimentally how common deep CNN architectures and Deep Learning practices can be used to train classifiers that recognize damages in a specialized domain, with application-specific metrics. We present methods for synthesizing, pre-processing and re-sampling of the necessary dataset. It is shown that pre-trained, parameter-efficient CNN architectures that implement skip-connections, complemented by
global max-pooling before the final layer, are well suited for that dataset. The Xception architecture has been chosen as backbone for the classifier due to its high recall and fast convergence. To mitigate the detrimental influence of imbalanced training data, training data re-sampling that equalizes the class distribution is implemented. It has a positive effect recall, especially
when applied to multi-label classification. When using re-sampling and data augmentation, the performance of multi-label and single-label classification can be brought to the same level. However, the best achieved FNR is 5.4%, with a softmax classifier, combining all regularization methods.
Finally, we investigated how regularization can be used to increase generalization capability with limited training data. Data augmentation is the most effective regularization method, even though its full potential has not been explored yet. Dropout benefits single-label classification
but not multi-label classification. L2-regularization has a moderate positive effect on both. Naively combining the regularization techniques without an exhaustive grid search or automated search on average does not yield any additional gains and shows the limit of manual hyper-parameter tuning.
feasibility of using a deep Convolutional Neural Network (CNN) for (semi-)automated damage detection, with the goal of achieving a high recall (low False Negative Rate (FNR)) on the small damages. The problem is framed as a classification problem on limited and imbalanced data. However, it is not pre-defined if single-label or multi-label classification should be used,
and both approaches are investigated.
The main contribution of this work is to show experimentally how common deep CNN architectures and Deep Learning practices can be used to train classifiers that recognize damages in a specialized domain, with application-specific metrics. We present methods for synthesizing, pre-processing and re-sampling of the necessary dataset. It is shown that pre-trained, parameter-efficient CNN architectures that implement skip-connections, complemented by
global max-pooling before the final layer, are well suited for that dataset. The Xception architecture has been chosen as backbone for the classifier due to its high recall and fast convergence. To mitigate the detrimental influence of imbalanced training data, training data re-sampling that equalizes the class distribution is implemented. It has a positive effect recall, especially
when applied to multi-label classification. When using re-sampling and data augmentation, the performance of multi-label and single-label classification can be brought to the same level. However, the best achieved FNR is 5.4%, with a softmax classifier, combining all regularization methods.
Finally, we investigated how regularization can be used to increase generalization capability with limited training data. Data augmentation is the most effective regularization method, even though its full potential has not been explored yet. Dropout benefits single-label classification
but not multi-label classification. L2-regularization has a moderate positive effect on both. Naively combining the regularization techniques without an exhaustive grid search or automated search on average does not yield any additional gains and shows the limit of manual hyper-parameter tuning.
This work explores the possibility of incorporating depth information into a deep neural network to improve accuracy of RGB instance segmentation. The baseline of this work is semantic instance segmentation with discriminative loss function.The baseline work proposes a novel discriminative loss function with which the semantic net-work can learn a n-D embedding for all pixels belonging to instances. Embeddings of the same instances are attracted to their own centers while centers of different instance embeddings repulse each other. Two limitations are set for attraction and repulsion, namely the in-margin and out-margin. A post-processing procedure (clustering) is required to infer instance indices from embeddings with an important parameter bandwidth, the threshold for clustering. The contribution of the work in this thesis are several new methods to incorporate depth information into the baseline work. One simple method is adding scaled depth directly to RGB embeddings, which is named as scaling. Through theorizing and experiments, this work also proposes that depth pixels can be encoded into 1-D embeddings with the same discriminative loss function and combined with RGB embeddings. Explored combination methods are fusion and concatenation. Additionally, two depth pre-processing methods are proposed, replication and coloring. From the experimental result, both scaling and fusion lead to significant improvements over baseline work while concatenation contributes more to classes with lots of similarities.
...
This work explores the possibility of incorporating depth information into a deep neural network to improve accuracy of RGB instance segmentation. The baseline of this work is semantic instance segmentation with discriminative loss function.The baseline work proposes a novel discriminative loss function with which the semantic net-work can learn a n-D embedding for all pixels belonging to instances. Embeddings of the same instances are attracted to their own centers while centers of different instance embeddings repulse each other. Two limitations are set for attraction and repulsion, namely the in-margin and out-margin. A post-processing procedure (clustering) is required to infer instance indices from embeddings with an important parameter bandwidth, the threshold for clustering. The contribution of the work in this thesis are several new methods to incorporate depth information into the baseline work. One simple method is adding scaled depth directly to RGB embeddings, which is named as scaling. Through theorizing and experiments, this work also proposes that depth pixels can be encoded into 1-D embeddings with the same discriminative loss function and combined with RGB embeddings. Explored combination methods are fusion and concatenation. Additionally, two depth pre-processing methods are proposed, replication and coloring. From the experimental result, both scaling and fusion lead to significant improvements over baseline work while concatenation contributes more to classes with lots of similarities.