Effective Mel-Spectrogram Synthesis for Multi-Speaker Text-to-Speech with Band-Specific Acoustic Modeling

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Recent advances in multi-speaker text-to-speech synthesis have enabled speaker-adaptive generation from short reference samples, yet preserving speaker individuality while sustaining real-time inference remains difficult. Conventional architectures process the entire Mel spectrogram with shared convolutional filters, leading to low-frequency energy patterns dominating and insufficient modeling of high-frequency cues that contain fine-grained spectral variations. To address this imbalance, this paper introduces a frequency-aware architecture that allocates dedicated modeling capacity to regions with distinct acoustic characteristics while keeping the overall computation lightweight. The high-frequency pathway focuses on rapid spectral transitions and localized acoustic variations, whereas the low-frequency pathway emphasizes global structure and rhythmic stability. A second-stage refinement module integrates these complementary representations through selective feature calibration with negligible overhead. Experiments on the LibriTTS dataset demonstrate that the proposed method achieves improved spectral accuracy, reducing Mel-cepstral distortion to 10.384, while maintaining near real-time inference with a real-time factor of 0.178.

키워드

Multi-Speaker Text-To-SpeechRepresentation LearningSpeech Synthesis
제목
Effective Mel-Spectrogram Synthesis for Multi-Speaker Text-to-Speech with Band-Specific Acoustic Modeling
저자
Moon, A-SeongSong, DahyunLee, Jaesung
DOI
10.1109/ACCESS.2026.3706694
발행일
2026
유형
Article
저널명
IEEE Access
14
페이지
101942 ~ 101954