Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Nov 26, 2024

Wei Guo, Heng Wang, Jianbo Ma, Weidong Cai

Figure 1 for Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Figure 2 for Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Figure 3 for Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Figure 4 for Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Share this with someone who'll enjoy it:

Abstract:Vision-to-audio (V2A) synthesis has broad applications in multimedia. Recent advancements of V2A methods have made it possible to generate relevant audios from inputs of videos or still images. However, the immersiveness and expressiveness of the generation are limited. One possible problem is that existing methods solely rely on the global scene and overlook details of local sounding objects (i.e., sound sources). To address this issue, we propose a Sound Source-Aware V2A (SSV2A) generator. SSV2A is able to locally perceive multimodal sound sources from a scene with visual detection and cross-modality translation. It then contrastively learns a Cross-Modal Sound Source (CMSS) Manifold to semantically disambiguate each source. Finally, we attentively mix their CMSS semantics into a rich audio representation, from which a pretrained audio generator outputs the sound. To model the CMSS manifold, we curate a novel single-sound-source visual-audio dataset VGGS3 from VGGSound. We also design a Sound Source Matching Score to measure localized audio relevance. This is to our knowledge the first work to address V2A generation at the sound-source level. Extensive experiments show that SSV2A surpasses state-of-the-art methods in both generation fidelity and relevance. We further demonstrate SSV2A's ability to achieve intuitive V2A control by compositing vision, text, and audio conditions. Our SSV2A generation can be tried and heard at https://ssv2a.github.io/SSV2A-demo .

* 16 pages, 9 figures, source code released at https://github.com/wguo86/SSV2A

View paper on

Share this with someone who'll enjoy it:

Title:Gotta Hear Them All: Sound Source Aware Vision to Audio Generation

Paper and Code