상세 보기
Effective Mel-Spectrogram Synthesis for Multi-Speaker Text-to-Speech with Band-Specific Acoustic Modeling
- Moon, A-Seong;
- Song, Dahyun;
- Lee, Jaesung
WEB OF SCIENCE
0SCOPUS
0초록
Recent advances in multi-speaker text-to-speech synthesis have enabled speaker-adaptive generation from short reference samples, yet preserving speaker individuality while sustaining real-time inference remains difficult. Conventional architectures process the entire Mel spectrogram with shared convolutional filters, leading to low-frequency energy patterns dominating and insufficient modeling of high-frequency cues that contain fine-grained spectral variations. To address this imbalance, this paper introduces a frequency-aware architecture that allocates dedicated modeling capacity to regions with distinct acoustic characteristics while keeping the overall computation lightweight. The high-frequency pathway focuses on rapid spectral transitions and localized acoustic variations, whereas the low-frequency pathway emphasizes global structure and rhythmic stability. A second-stage refinement module integrates these complementary representations through selective feature calibration with negligible overhead. Experiments on the LibriTTS dataset demonstrate that the proposed method achieves improved spectral accuracy, reducing Mel-cepstral distortion to 10.384, while maintaining near real-time inference with a real-time factor of 0.178.
키워드
- 제목
- Effective Mel-Spectrogram Synthesis for Multi-Speaker Text-to-Speech with Band-Specific Acoustic Modeling
- 저자
- Moon, A-Seong; Song, Dahyun; Lee, Jaesung
- 발행일
- 2026
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 101942 ~ 101954