Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Kanami Imamura

Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals

Jun 25, 2024

Kentaro Seki, Shinnosuke Takamichi, Norihiro Takamune, Yuki Saito, Kanami Imamura, Hiroshi Saruwatari

Figure 1 for Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals

Figure 2 for Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals

Figure 3 for Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals

Figure 4 for Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals

Abstract:This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms, ignoring the stereo listening experience inherent in human hearing. Our baseline approach addresses this gap by integrating blind source separation (BSS), voice conversion (VC), and spatial mixing to handle multi-channel waveforms. Through experimental evaluations, we organize and identify the key challenges inherent in this task, such as maintaining audio quality and accurately preserving spatial information. Our results highlight the fundamental difficulties in balancing these aspects, providing a benchmark for future research in spatial voice conversion. The proposed method's code is publicly available to encourage further exploration in this domain.

* Accepted to Interspeech 2024

Via

Access Paper or Ask Questions

Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides

Jun 19, 2023

Kanami Imamura, Tomohiko Nakamura, Norihiro Takamune, Kohei Yatabe, Hiroshi Saruwatari

Figure 1 for Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides

Figure 2 for Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides

Figure 3 for Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides

Figure 4 for Algorithms of Sampling-Frequency-Independent Layers for Non-integer Strides

Abstract:In this paper, we propose algorithms for handling non-integer strides in sampling-frequency-independent (SFI) convolutional and transposed convolutional layers. The SFI layers have been developed for handling various sampling frequencies (SFs) by a single neural network. They are replaceable with their non-SFI counterparts and can be introduced into various network architectures. However, they could not handle some specific configurations when combined with non-SFI layers. For example, an SFI extension of Conv-TasNet, a standard audio source separation model, cannot handle some pairs of trained and target SFs because the strides of the SFI layers become non-integers. This problem cannot be solved by simple rounding or signal resampling, resulting in the significant performance degradation. To overcome this problem, we propose algorithms for handling non-integer strides by using windowed sinc interpolation. The proposed algorithms realize the continuous-time representations of features using the interpolation and enable us to sample instants with the desired stride. Experimental results on music source separation showed that the proposed algorithms outperformed the rounding- and signal-resampling-based methods at SFs lower than the trained SF.

* 5 pages, 3 figures, accepted for European Signal Processing Conference 2023 (EUSIPCO 2023)

Via

Access Paper or Ask Questions