상세 보기
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
- Park, Jisoo;
- Lee, Seonghak;
- Kim, Guisik;
- Kim, Taewoo;
- Kwon, Junseok
WEB OF SCIENCE
0SCOPUS
0초록
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivating the need for a unified solution. While recent approaches have attempted to integrate SE and SS within multi-stage architectures, these approaches typically involve complex, parameter-heavy models and rely on supervised training, limiting scalability and generalization. In this work, we propose UniVoiceLite, a lightweight and unsupervised audiovisual framework that unifies SE and SS within a single model. UniVoiceLite leverages lip motion and facial identity cues to guide speech extraction and employs Wasserstein distance regularization to stabilize the latent space without requiring paired noisy-clean data. Experimental results demonstrate that UniVoiceLite achieves strong performance in both noisy and multi-speaker scenarios, combining efficiency with robust generalization. The source code is available at https://github.com/jisoo-o/UniVoiceLite.
키워드
- 제목
- Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
- 저자
- Park, Jisoo; Lee, Seonghak; Kim, Guisik; Kim, Taewoo; Kwon, Junseok
- 발행일
- 2025
- 유형
- Proceedings Paper
- 저널명
- ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
- 언어
- ENG
- 출판사
- Institute of Electrical and Electronics Engineers Inc.
- ISSN
- P 2997-6928