Using various pre-trained models for audio feature extraction in automated audio captioning

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

5

초록

The DCASE automated audio captioning challenge aimed to construct a model that generates captions describing given audio. Our team developed a CNN14 encoder (pre-trained on AudioSet data) along with a Transformer decoder model that ranked sixth place in the competition. Many teams utilized pre-trained networks, and it was evident that more research into their utilization was required. This paper presented comprehensive experiments conducted with various encoder networks for the proposed system, including CNN10, CNN14 ResNet54, AST, VGGNet, and EfficientNet. The pre-trained networks of CNN10, CNN14, ResNet54, and AST were trained on AudioSet data, while the pre-trained networks of AST, VGGNet, and EfficientNet were trained on ImageNet data. The best outcomes were achieved when the pre-trained CNN10, trained on AudioSet data, was utilized as an encoder with the Transformer serving as a decoder, and fine-tuning applied. Moreover, a qualitative study confirmed that our model generates plausible captions for different types of audio. © 2023 Elsevier Ltd

키워드

Acoustic scene detection; Audio captioning; Convolutional neural network; Encoder–decoder; Transfer learning; Transformer
제목
Using various pre-trained models for audio feature extraction in automated audio captioning
저자
Won, Hyejin; Kim, Baekseung; Kwak, Il-Youp; Lim, Changwon
DOI
10.1016/j.eswa.2023.120664
발행일
2023-11
유형
Article
저널명
Expert Systems with Applications
권
231