Resmax: Detecting voice spoofing attacks with residual network and max feature map

Citations

WEB OF SCIENCE

14
Citations

SCOPUS

24

초록

The “2019 Automatic Speaker Verification Spoofing And Countermeasures Challenge” (ASVspoof) competition aimed to facilitate the design of highly accurate voice spoofing attack detection systems. the competition did not emphasize model complexity and latency requirements; such constraints are strict and integral in real-world deployment. Hence, most of the top performing solutions from the competition all used an ensemble approach, and combined multiple complex deep learning models to maximize detection accuracy - this kind of approach would sit uneasily with real-world deployment constraints. To design a lightweight system, we combined the notions of skip connection (from ResNet) and max feature map (from Light CNN), and evaluated the accuracy of the system using the ASVspoof 2019 dataset. With an optimized constant Q transform (CQT) feature, our single model achieved a replay attack detection equal error rate (EER) of 0.37% on the evaluation set, surpassing the top ensemble system from the competition that achieved an EER of 0.39%. © 2020 IEEE

키워드

Voice assistant security; Voice presentation attack detection; Voice spoofing attack; Voice synthesis attack; Complex networks; Deep learning; Feature extraction; Automatic speaker verification; Constant q transforms; Detection accuracy; Ensemble approaches; Lightweight systems; Model complexity; Real world deployment; Spoofing attacks; Speech recognition
제목
Resmax: Detecting voice spoofing attacks with residual network and max feature map
저자
Kwak, I.-Y.; Kwag, S.; Lee, J.; Huh, J.H.; Lee, C.-H.; Jeon, Y.; Hwang, J.; Yoon, J.W.
DOI
10.1109/ICPR48806.2021.9412165
발행일
2021
유형
Proceedings Paper
저널명
Proceedings - International Conference on Pattern Recognition
페이지
4837 ~ 4844