상세 보기
Diffusion Model-Based Generative Pipeline for Children Song Video
- Lee, Sanghyuck;
- Khairulov, Timur;
- Lee, Jaesung
SCOPUS
0초록
Children songs have been essential in early childhood education, supporting cognitive development, language acquisition, and emotional expression. With the rise of digital media, the traditional children songs have evolved into multimedia experiences, including music videos. However, the creation of these videos is a resource-intensive process that requires a blend of artistic and technical expertise. Meanwhile, recent advancements in generative models, especially diffusion models, have shown impressive text-to-image capabilities, though they still face limitations in generating temporally coherent video content. This paper explores an innovative approach to generating music videos for children songs, that convert children song lyrics into visually appealing and contextually relevant video content. Our approach integrates natural language processing to interpret lyrics and computer vision techniques to generate corresponding animations and visuals. Our experiments on 20 prompts demonstrate that the Cascade SD model outperforms the other four models across three evaluation measures. The qualitative analysis on 10 prompts demonstrates the superiority of the Cascade SD model and highlights the effectiveness of negative prompting and secondary prompting techniques. The demo is available in https://github.com/tkdgur658/Children_Song_Video. © 2025 IEEE.
키워드
- 제목
- Diffusion Model-Based Generative Pipeline for Children Song Video
- 저자
- Lee, Sanghyuck; Khairulov, Timur; Lee, Jaesung
- 발행일
- 2025
- 유형
- Conference paper
- 저널명
- Digest of Technical Papers - IEEE International Conference on Consumer Electronics
- 언어
- ENG
- 출판사
- Institute of Electrical and Electronics Engineers Inc.
- ISSN
- P 0747-668X