Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

May 03, 2021

Yan-Bo Lin, Yu-Chiang Frank Wang

Figure 1 for Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

Figure 2 for Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

Figure 3 for Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

Figure 4 for Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

Share this with someone who'll enjoy it:

Abstract:Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient information. To address this issue, we propose an audio spatialization framework to convert a monaural video into a binaural one exploiting the relationship across audio and visual components. By preserving the left-right consistency in both audio and visual modalities, our learning strategy can be viewed as a self-supervised learning technique, and alleviates the dependency on a large amount of video data with ground truth binaural audio data during training. Experiments on benchmark datasets confirm the effectiveness of our proposed framework in both semi-supervised and fully supervised scenarios, with ablation studies and visualization further support the use of our model for audio spatialization.

* AAAI'21

View paper on

Share this with someone who'll enjoy it:

Title:Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

Paper and Code