No frame left behind

None, None; None, None; None, None; None, None; None, None

No frame left behind

Full Video Action Recognition

Conference Paper (2021)

Author(s)

Xin Liu (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Silvia L. Pintea (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Fatemeh Karimi Nejadasl (TomTom BV)

Olaf Booij (TomTom BV)

Jan C. van Gemert (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group

Pattern Recognition and Bioinformatics

DOI related publication

https://doi.org/10.1109/CVPR46437.2021.01465 Final published version

To reference this document use

https://resolver.tudelft.nl/uuid:36a98ed0-ad25-4691-a6c3-98ac48311104

More Info

expand_more

Publication Year

2021

Language

English

Research Group

Pattern Recognition and Bioinformatics

Article number

9578276

Pages (from-to)

14887-14896

ISBN (print)

978-1-6654-4510-8

ISBN (electronic)

978-1-6654-4509-2

Event

2021 IEEE/CVF Conference on Computer Vision<br/>and Pattern Recognition (2021-06-20 - 2021-06-25), Virtual at Nashville, United States

Downloads counter

205

Abstract

Not all video frames are equally informative for recognizing an action. It is computationally infeasible to train deep networks on all video frames when actions develop over hundreds of frames. A common heuristic is uniformly sampling a small number of video frames and using these to recognize the action. Instead, here we propose full video action recognition and consider all video frames. To make this computational tractable, we first cluster all frame activations along the temporal dimension based on their similarity with respect to the classification task, and then temporally aggregate the frames in the clusters into a smaller number of representations. Our method is end-to-end trainable and computationally efficient as it relies on temporally localized clustering in combination with fast Hamming distances in feature space. We evaluate on UCF101, HMDB51, Breakfast, and Something-Something V1 and V2, where we compare favorably to existing heuristic frame sampling methods.