Picture for Chaoren Wang

Chaoren Wang

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Add code
Jun 30, 2026
Viaarxiv icon

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora

Add code
Apr 13, 2026
Viaarxiv icon

Aliasing-Free Neural Audio Synthesis

Add code
Dec 23, 2025
Figure 1 for Aliasing-Free Neural Audio Synthesis
Figure 2 for Aliasing-Free Neural Audio Synthesis
Figure 3 for Aliasing-Free Neural Audio Synthesis
Figure 4 for Aliasing-Free Neural Audio Synthesis
Viaarxiv icon

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

Add code
Nov 11, 2025
Figure 1 for SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
Figure 2 for SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
Figure 3 for SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
Figure 4 for SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
Viaarxiv icon

SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level

Add code
Oct 30, 2025
Viaarxiv icon

Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN

Add code
May 21, 2025
Viaarxiv icon

DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation

Add code
May 19, 2025
Viaarxiv icon

SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset

Add code
May 14, 2025
Viaarxiv icon

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

Add code
May 07, 2025
Viaarxiv icon

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Add code
Jan 27, 2025
Viaarxiv icon