Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Move2Hear: Active Audio-Visual Source Separation

May 15, 2021

Sagnik Majumder, Ziad Al-Halah, Kristen Grauman

Figure 1 for Move2Hear: Active Audio-Visual Source Separation

Figure 2 for Move2Hear: Active Audio-Visual Source Separation

Figure 3 for Move2Hear: Active Audio-Visual Source Separation

Figure 4 for Move2Hear: Active Audio-Visual Source Separation

Share this with someone who'll enjoy it:

Abstract:We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources simultaneously (e.g., a person speaking down the hall in a noisy household) and must use its eyes and ears to automatically separate out the sounds originating from the target object within a limited time budget. Towards this goal, we introduce a reinforcement learning approach that trains movement policies controlling the agent's camera and microphone placement over time, guided by the improvement in predicted audio separation quality. We demonstrate our approach in scenarios motivated by both augmented reality (system is already co-located with the target object) and mobile robotics (agent begins arbitrarily far from the target object). Using state-of-the-art realistic audio-visual simulations in 3D environments, we demonstrate our model's ability to find minimal movement sequences with maximal payoff for audio source separation. Project: http://vision.cs.utexas.edu/projects/move2hear.

View paper on

Share this with someone who'll enjoy it:

Title:Move2Hear: Active Audio-Visual Source Separation

Paper and Code