Picture for Yuancheng Wang

Yuancheng Wang

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Add code
Feb 05, 2025
Viaarxiv icon

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Add code
Jan 27, 2025
Viaarxiv icon

Overview of the Amphion Toolkit (v0.2)

Add code
Jan 26, 2025
Viaarxiv icon

AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

Add code
Jan 26, 2025
Viaarxiv icon

Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities

Add code
Nov 29, 2024
Figure 1 for Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities
Figure 2 for Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities
Figure 3 for Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities
Figure 4 for Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities
Viaarxiv icon

Debatts: Zero-Shot Debating Text-to-Speech Synthesis

Add code
Nov 10, 2024
Figure 1 for Debatts: Zero-Shot Debating Text-to-Speech Synthesis
Figure 2 for Debatts: Zero-Shot Debating Text-to-Speech Synthesis
Figure 3 for Debatts: Zero-Shot Debating Text-to-Speech Synthesis
Figure 4 for Debatts: Zero-Shot Debating Text-to-Speech Synthesis
Viaarxiv icon

MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Add code
Sep 01, 2024
Viaarxiv icon

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Add code
Jul 07, 2024
Figure 1 for Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Figure 2 for Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Figure 3 for Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Figure 4 for Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Viaarxiv icon

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Add code
Jul 01, 2024
Figure 1 for FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Figure 2 for FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Figure 3 for FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Figure 4 for FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Viaarxiv icon

SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Add code
Jun 19, 2024
Viaarxiv icon