Diffusion Model-Based Generative Pipeline for Children Song Video

Citations

SCOPUS

0

초록

Children songs have been essential in early childhood education, supporting cognitive development, language acquisition, and emotional expression. With the rise of digital media, the traditional children songs have evolved into multimedia experiences, including music videos. However, the creation of these videos is a resource-intensive process that requires a blend of artistic and technical expertise. Meanwhile, recent advancements in generative models, especially diffusion models, have shown impressive text-to-image capabilities, though they still face limitations in generating temporally coherent video content. This paper explores an innovative approach to generating music videos for children songs, that convert children song lyrics into visually appealing and contextually relevant video content. Our approach integrates natural language processing to interpret lyrics and computer vision techniques to generate corresponding animations and visuals. Our experiments on 20 prompts demonstrate that the Cascade SD model outperforms the other four models across three evaluation measures. The qualitative analysis on 10 prompts demonstrates the superiority of the Cascade SD model and highlights the effectiveness of negative prompting and secondary prompting techniques. The demo is available in https://github.com/tkdgur658/Children_Song_Video. © 2025 IEEE.

키워드

generative AI; music video generation; text-to-video generation
제목
Diffusion Model-Based Generative Pipeline for Children Song Video
저자
Lee, Sanghyuck; Khairulov, Timur; Lee, Jaesung
DOI
10.1109/ICCE63647.2025.10930103
발행일
2025
유형
Conference paper
저널명
Digest of Technical Papers - IEEE International Conference on Consumer Electronics