Towards Decoding Motion: How Depth and Optical Flow Inform Neural Ego-Motion Estimation
Q. Missinne (TU Delft - Mechanical Engineering)
J.F.P. Kooij – Mentor (TU Delft - Mechanical Engineering)
G.C.H.E. de Croon – Mentor (TU Delft - Aerospace Engineering)
M. Zaffar – Mentor (TU Delft - Mechanical Engineering)
S.A. Bahnam – Mentor (TU Delft - Aerospace Engineering)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Deep neural networks are increasingly used for ego-motion estimation. Often, in self-supervised ego-motion networks, it is decoded from a depth network and until now has not been decoded from an optical flow network. This is surprising given the tight relationship between optical flow and ego-motion. While both representations are widely used in learning-based approaches, the extent to which their latent space encodes motion information remains poorly understood. This paper presents a controlled analysis of how well depth-based and optical flow-based neural networks encode ego-motion. Using supervised depth, flow and pose networks trained on TartanAirV2, we probe motion information by attaching identical minimal pose decoders to frozen encoders. Representational space analysis through Centered Kernel Analysis (CKA) and feature space analysis through Principal Component Analysis (PCA) are used to examine how motion information is structured across network hierarchies. This work shows that optical flow representations encode ego-motion information more explicitly, in a lower dimensional, linearly accessible structure as opposed to depth representations, which exhibit weak alignment with pose. These findings suggest that in a self-supervised setting, ego-motion estimation can best be decoded from an optical flow network as opposed to a depth network.