상세 보기
RawNet3 화자 표현을 활용한 임의의 화자 간음성 변환을 위한 StarGAN의 확장
- 박보경;
- 박소민;
- 홍현기
초록
Voice conversion, a technology that allows an individual’s speech data to be regenerated with the acoustic properties(tone, cadence,gender) of another, has countless applications in education, communication, and entertainment. This paper proposes an approach basedon the StarGAN-VC model that generates realistic-sounding speech without requiring parallel utterances. To overcome the constraintsof the existing StarGAN-VC model that utilizes one-hot vectors of original and target speaker information, this paper extracts featurevectors of target speakers using a pre-trained version of Rawnet3. This results in a latent space where voice conversion can be performedwithout direct speaker-to-speaker mappings, enabling an any-to-any structure. In addition to the loss terms used in the originalStarGAN-VC model, Wasserstein distance is used as a loss term to ensure that generated voice segments match the acoustic propertiesof the target voice. Two Time-Scale Update Rule (TTUR) is also used to facilitate stable training. Experimental results show that theproposed method outperforms previous methods, including the StarGAN-VC network on which it was based.
키워드
- 제목
- RawNet3 화자 표현을 활용한 임의의 화자 간음성 변환을 위한 StarGAN의 확장
- 제목 (타언어)
- Extending StarGAN-VC to Unseen Speakers Using RawNet3 Speaker Representation
- 저자
- 박보경; 박소민; 홍현기
- 발행일
- 2023-07
- 권
- 12
- 호
- 7
- 페이지
- 303 ~ 314