Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Srdan Kitic

SALADnet: Self-Attentive multisource Localization in the Ambisonics Domain

Jul 23, 2021

Pierre-Amaury Grumiaux, Srdan Kitic, Prerak Srivastava, Laurent Girin, Alexandre Guérin

Figure 1 for SALADnet: Self-Attentive multisource Localization in the Ambisonics Domain

Figure 2 for SALADnet: Self-Attentive multisource Localization in the Ambisonics Domain

Figure 3 for SALADnet: Self-Attentive multisource Localization in the Ambisonics Domain

Figure 4 for SALADnet: Self-Attentive multisource Localization in the Ambisonics Domain

Abstract:In this work, we propose a novel self-attention based neural network for robust multi-speaker localization from Ambisonics recordings. Starting from a state-of-the-art convolutional recurrent neural network, we investigate the benefit of replacing the recurrent layers by self-attention encoders, inherited from the Transformer architecture. We evaluate these models on synthetic and real-world data, with up to 3 simultaneous speakers. The obtained results indicate that the majority of the proposed architectures either perform on par, or outperform the CRNN baseline, especially in the multisource scenario. Moreover, by avoiding the recurrent layers, the proposed models lend themselves to parallel computing, which is shown to produce considerable savings in execution time.

* Accepted to Workshop on Applications of Signal Processing to Audio and Acoustics

Via

Access Paper or Ask Questions

Improved feature extraction for CRNN-based multiple sound source localization

May 05, 2021

Pierre-Amaury Grumiaux, Srdan Kitic, Laurent Girin, Alexandre Guérin

Figure 1 for Improved feature extraction for CRNN-based multiple sound source localization

Figure 2 for Improved feature extraction for CRNN-based multiple sound source localization

Figure 3 for Improved feature extraction for CRNN-based multiple sound source localization

Figure 4 for Improved feature extraction for CRNN-based multiple sound source localization

Abstract:In this work, we propose to extend a state-of-the-art multi-source localization system based on a convolutional recurrent neural network and Ambisonics signals. We significantly improve the performance of the baseline network by changing the layout between convolutional and pooling layers. We propose several configurations with more convolutional layers and smaller pooling sizes in-between, so that less information is lost across the layers, leading to a better feature extraction. In parallel, we test the system's ability to localize up to 3 sources, in which case the improved feature extraction provides the most significant boost in accuracy. We evaluate and compare these improved configurations on synthetic and real-world data. The obtained results show a quite substantial improvement of the multiple sound source localization performance over the baseline network.

* 5 pages, 2 figures. Accepted to EUSIPCO 2021

Via

Access Paper or Ask Questions

Multichannel CRNN for Speaker Counting: an Analysis of Performance

Jan 06, 2021

Pierre-Amaury Grumiaux, Srdan Kitic, Laurent Girin, Alexandre Guérin

Figure 1 for Multichannel CRNN for Speaker Counting: an Analysis of Performance

Figure 2 for Multichannel CRNN for Speaker Counting: an Analysis of Performance

Abstract:Speaker counting is the task of estimating the number of people that are simultaneously speaking in an audio recording. For several audio processing tasks such as speaker diarization, separation, localization and tracking, knowing the number of speakers at each timestep is a prerequisite, or at least it can be a strong advantage, in addition to enabling a low latency processing. In a previous work, we addressed the speaker counting problem with a multichannel convolutional recurrent neural network which produces an estimation at a short-term frame resolution. In this work, we show that, for a given frame, there is an optimal position in the input sequence for best prediction accuracy. We empirically demonstrate the link between that optimal position, the length of the input sequence and the size of the convolutional filters.

* Presented at Forum Acusticum 2020

Via

Access Paper or Ask Questions

Unifying local and non-local signal processing with graph CNNs

Jul 07, 2017

Gilles Puy, Srdan Kitic, Patrick Pérez

Figure 1 for Unifying local and non-local signal processing with graph CNNs

Figure 2 for Unifying local and non-local signal processing with graph CNNs

Figure 3 for Unifying local and non-local signal processing with graph CNNs

Figure 4 for Unifying local and non-local signal processing with graph CNNs

Abstract:This paper deals with the unification of local and non-local signal processing on graphs within a single convolutional neural network (CNN) framework. Building upon recent works on graph CNNs, we propose to use convolutional layers that take as inputs two variables, a signal and a graph, allowing the network to adapt to changes in the graph structure. In this article, we explain how this framework allows us to design a novel method to perform style transfer.

Via

Access Paper or Ask Questions