Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Dec 09, 2024

Tianxin Xie, Yan Rong, Pengfei Zhang, Li Liu

Figure 1 for Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Figure 2 for Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Figure 3 for Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Figure 4 for Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Share this with someone who'll enjoy it:

Abstract:Text-to-speech (TTS), also known as speech synthesis, is a prominent research area that aims to generate natural-sounding human speech from text. Recently, with the increasing industrial demand, TTS technologies have evolved beyond synthesizing human-like speech to enabling controllable speech generation. This includes fine-grained control over various attributes of synthesized speech such as emotion, prosody, timbre, and duration. Besides, advancements in deep learning, such as diffusion and large language models, have significantly enhanced controllable TTS over the past several years. In this paper, we conduct a comprehensive survey of controllable TTS, covering approaches ranging from basic control techniques to methods utilizing natural language prompts, aiming to provide a clear understanding of the current state of research. We examine the general controllable TTS pipeline, challenges, model architectures, and control strategies, offering a comprehensive and clear taxonomy of existing methods. Additionally, we provide a detailed summary of datasets and evaluation metrics and shed some light on the applications and future directions of controllable TTS. To the best of our knowledge, this survey paper provides the first comprehensive review of emerging controllable TTS methods, which can serve as a beneficial resource for both academic researchers and industry practitioners.

* A comprehensive survey on controllable TTS, 23 pages, 6 tables, 4 figures, 280 references

View paper on

Share this with someone who'll enjoy it:

Title:Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

Paper and Code