RawNet3 화자 표현을 활용한 임의의 화자 간음성 변환을 위한 StarGAN의 확장

Extending StarGAN-VC to Unseen Speakers Using RawNet3 Speaker Representation

초록

Voice conversion, a technology that allows an individual’s speech data to be regenerated with the acoustic properties(tone, cadence,gender) of another, has countless applications in education, communication, and entertainment. This paper proposes an approach basedon the StarGAN-VC model that generates realistic-sounding speech without requiring parallel utterances. To overcome the constraintsof the existing StarGAN-VC model that utilizes one-hot vectors of original and target speaker information, this paper extracts featurevectors of target speakers using a pre-trained version of Rawnet3. This results in a latent space where voice conversion can be performedwithout direct speaker-to-speaker mappings, enabling an any-to-any structure. In addition to the loss terms used in the originalStarGAN-VC model, Wasserstein distance is used as a loss term to ensure that generated voice segments match the acoustic propertiesof the target voice. Two Time-Scale Update Rule (TTUR) is also used to facilitate stable training. Experimental results show that theproposed method outperforms previous methods, including the StarGAN-VC network on which it was based.

키워드

음성 변환화자 특성일반화StarGAN-VCRawNet3Voice ConversionSpeaker AttributeGeneralizationStarGAN-VCRawNet3
제목
RawNet3 화자 표현을 활용한 임의의 화자 간음성 변환을 위한 StarGAN의 확장
제목 (타언어)
Extending StarGAN-VC to Unseen Speakers Using RawNet3 Speaker Representation
저자
박보경박소민홍현기
발행일
2023-07
저널명
정보처리학회논문지. 소프트웨어 및 데이터 공학
12
7
페이지
303 ~ 314