Briefing of Text-To-Speech Trends Toward Zero-Shot Multi-Speaker Synthesis

Citations

SCOPUS

0

초록

This paper reviews advancements in Text-to-Speech (TTS) systems and focuses on the evolution from single-speaker systems to Zero-Shot Multi-Speaker TTS (ZS-TTS) systems. Traditional single-speaker TTS models require retraining whenever a new speaker is introduced, which limits scalability. Multi-speaker models address this issue but rely on large amounts of data and computational resources. ZS-TTS is a recent innovation that synthesizes natural speech for unseen speakers using only a few seconds of reference audio and eliminates the need for additional fine-tuning. In this paper, we introduce important methods such as speaker embedding, GAN-based models, and transformers, explaining how these approaches improve efficiency and voice quality. It also examines the challenges of inference speed and model scalability, especially in edge computing environments, and suggests potential directions for future research in TTS. © 2025 IEEE.

키워드

Speech Synthesis; Text-to-Speech; Zero-Shot Multi-Speaker TTS
제목
Briefing of Text-To-Speech Trends Toward Zero-Shot Multi-Speaker Synthesis
저자
Song, Dahyun; Lee, Jaesung
DOI
10.1109/ICCE63647.2025.10930170
발행일
2025
유형
Conference paper
저널명
Digest of Technical Papers - IEEE International Conference on Consumer Electronics