SpeechCAT

Cross-Attentive Transformer for Audio to Motion Generation

Conference Paper (2025)
Author(s)

Sebastian Deaconu (Student TU Delft)

Xiangwei Shi (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Thomas Markhorst (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Jouh Yeong Chew (Honda Research Institute Japan Co., Ltd.)

Xucong Zhang (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Pattern Recognition and Bioinformatics
DOI related publication
https://doi.org/10.1109/HRI61500.2025.10974020 Final published version
More Info
expand_more
Publication Year
2025
Language
English
Research Group
Pattern Recognition and Bioinformatics
Pages (from-to)
1284-1288
Publisher
IEEE
ISBN (electronic)
9798350378931
Event
20th Annual ACM/IEEE International Conference on Human-Robot Interaction, HRI 2025 (2025-03-04 - 2025-03-06), Melbourne, Australia
Downloads counter
28
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Audio-to-motion generation is an important task with applications in virtual avatar creation for XR systems and intelligent robot control in daily life scenarios. However, most existing motion generation methods rely on a single encoder-decoder architecture to model all body parts simultaneously, which limits their ability to capture the diverse and complex motions exhibited by humans. In this paper, we propose a novel method, SpeechCAT, that employs three separate encoder-decoder modules to individually model the motions of the face, body, and hands. To capture the relationships and synchronization among these body parts, we introduce a cross-attention mechanism to effectively learn their correlations. SpeechCAT ensures sufficient capacity to model the unique characteristics of each body part while preserving the coherence between them. Our experimental results demonstrate the superiority of SpeechCAT over baseline methods, highlighting its effectiveness in generating diverse, realistic, and synchronized motions with face, body, and hand parts.

Files

SpeechCAT_Cross-Attentive_Tran... (pdf)
(pdf | 1.28 Mb)
- Embargo expired in 30-10-2025
– Personal use only – Dutch Copyright Act (Article 25fa)