Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Teo Guichoux

Investigating the impact of 2D gesture representation on co-speech gesture generation

Jun 24, 2024

Teo Guichoux, Laure Soulier, Nicolas Obin, Catherine Pelachaud

Figure 1 for Investigating the impact of 2D gesture representation on co-speech gesture generation

Figure 2 for Investigating the impact of 2D gesture representation on co-speech gesture generation

Figure 3 for Investigating the impact of 2D gesture representation on co-speech gesture generation

Figure 4 for Investigating the impact of 2D gesture representation on co-speech gesture generation

Abstract:Co-speech gestures play a crucial role in the interactions between humans and embodied conversational agents (ECA). Recent deep learning methods enable the generation of realistic, natural co-speech gestures synchronized with speech, but such approaches require large amounts of training data. "In-the-wild" datasets, which compile videos from sources such as YouTube through human pose detection models, offer a solution by providing 2D skeleton sequences that are paired with speech. Concurrently, innovative lifting models have emerged, capable of transforming these 2D pose sequences into their 3D counterparts, leading to large and diverse datasets of 3D gestures. However, the derived 3D pose estimation is essentially a pseudo-ground truth, with the actual ground truth being the 2D motion data. This distinction raises questions about the impact of gesture representation dimensionality on the quality of generated motions, a topic that, to our knowledge, remains largely unexplored. In this work, we evaluate the impact of the dimensionality of the training data, 2D or 3D joint coordinates, on the performance of a multimodal speech-to-gesture deep generative model. We use a lifting model to convert 2D-generated sequences of body pose to 3D. Then, we compare the sequence of gestures generated directly in 3D to the gestures generated in 2D and lifted to 3D as post-processing.

* 8 pages. Paper accepted at WACAI 2024

Via

Access Paper or Ask Questions