PoolImagen: Text-To-Image Diffusion Models with an Efficient Transformer Without Attention

Citations

SCOPUS

0

초록

Recent advancements in the field of image generation models have been particularly notable for diffusion models. Imagen has the most remarkable image generation capabilities among these models, particularly at high resolutions. However, Imagen comes with limitations, as creating high-quality results requires considerable computational resources and lengthy training times. To address these limitations, we propose PoolImagen, a novel and improved variant of Imagen that combines high performance with low computational costs. PoolImagen introduces various improvements to overcome the constraints of Imagen. Notably, we adopted the idea, first propose in MetaFormer, which suggests replacing the attention module with a pooling structure in the transformer architecture of Imagen to mitigate the issues related to increased training costs and computational complexity. Additionally, considering the influence of text encoder size on text-To-image transformation quality, we incorporate the large language models (e.g. flan-T5-xxl), an extension of the t5 model that offers more parameters and refined text processing capabilities. With a well-Trained transformer, PoolImagen achieves image generation with consistent performance and significantly accelerated training velocities. In experiments based on bird image datasets, PoolImagen demonstrates improved performance in terms of Fréchet Inception Distance (FID) and training time. In the case of the bird datasets, PoolImagen exhibits an approximately 11.29% improvement in FID compared to Imagen, while training time is reduced by 2.25 times. In addition, we conducted additional experiments to evaluate ability of PoolImagen to represent domain-specific features in generated images. These findings emphasize the potential of PoolImagen as a powerful tool for rapidly generating text-To-image outputs and suggest promising directions for enhancing the future performance of diffusion models. © 2024 IEEE.

키워드

diffusion models; image generation; text-To-image synthesis; transformer
제목
PoolImagen: Text-To-Image Diffusion Models with an Efficient Transformer Without Attention
저자
Ku, Hyeeun; Lee, Minhyeok; Ko, Kanghyeok; Baek, Sun Jae
DOI
10.1109/ICAIIC60209.2024.10463464
발행일
2024-02
유형
Conference paper
저널명
6th International Conference on Artificial Intelligence in Information and Communication, ICAIIC 2024
페이지
123 ~ 128