Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Aug 30, 2023

Dongheon Lee, Jung-Woo Choi

Figure 1 for DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Figure 2 for DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Figure 3 for DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Figure 4 for DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Share this with someone who'll enjoy it:

Abstract:In this work, we present DeFTAN-II, an efficient multichannel speech enhancement model based on transformer architecture and subgroup processing. Despite the success of transformers in speech enhancement, they face challenges in capturing local relations, reducing the high computational complexity, and lowering memory usage. To address these limitations, we introduce subgroup processing in our model, combining subgroups of locally emphasized features with other subgroups containing original features. The subgroup processing is implemented in several blocks of the proposed network. In the proposed split dense blocks extracting spatial features, a pair of subgroups is sequentially concatenated and processed by convolution layers to effectively reduce the computational complexity and memory usage. For the F- and T-transformers extracting temporal and spectral relations, we introduce cross-attention between subgroups to identify relationships between locally emphasized and non-emphasized features. The dual-path feedforward network then aggregates attended features in terms of the gating of local features processed by dilated convolutions. Through extensive comparisons with state-of-the-art multichannel speech enhancement models, we demonstrate that DeFTAN-II with subgroup processing outperforms existing methods at significantly lower computational complexity. Moreover, we evaluate the model's generalization capability on real-world data without fine-tuning, which further demonstrates its effectiveness in practical scenarios.

* 13 pages, 6 figures, submitted to IEEE/ACM Trans. Audio, Speech, Lang. Process

View paper on

Share this with someone who'll enjoy it:

Title:DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing

Paper and Code