详细信息
RII-GAN: Multi-scaled Aligning-Based Reversed Image Interaction Network for Text-to-Image Synthesis ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:RII-GAN: Multi-scaled Aligning-Based Reversed Image Interaction Network for Text-to-Image Synthesis
作者:Yuan, Haofei[1];Zhu, Hongqing[1];Yang, Suyi[2];Wang, Ziying[1];Wang, Nan[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, 130 Meilong Rd, Shanghai 200237, Peoples R China;[2]Kings Coll London, Dept Math Nat Math & Engn Sci, London WC2R 2LS, England
年份:2024
卷号:56
期号:1
外文期刊名:NEURAL PROCESSING LETTERS
收录:;EI(收录号:20241115733479);WOS:【SCI-EXPANDED(收录号:WOS:001162756500005)】;
基金:The authors would like to thank the anonymous reviewers and the associate editor for their insightful comments that significantly improved the quality of this paper. This work was supported by the National Nature Science Foundation of China under Grant 61872143.
语种:英文
外文关键词:Text-to-image; Single-stage generation; Reversed image interaction network; Adaptive affine-based generator; Dual-channel
摘要:The text-to-image (T2I) model based on a single-stage generative adversarial network (GAN) has significantly succeeded in recent years. However, the generation model based on GAN has two disadvantages: the generator does not introduce any image feature manifold structure, which makes it challenging to align the image and text features. Another is the image's diversity; the text's abstraction will prevent the model from learning the actual image distribution. This paper proposes a reversed image interaction generative adversarial network (RII-GAN), which consists of four components: text encoder, reversed image interaction network (RIIN), adaptive affine-based generator, and dual-channel feature alignment discriminator (DFAD). RIIN indirectly introduces the actual image distribution into the generation network, thus overcoming the problem that the network lacks the learning of the actual image feature manifold structure and generating the distribution of text-matching images. Each adaptive affine block (AAB) in the proposed affine-based generator can adaptively enhance text information, establishing an updated relation between original independent fusion blocks and the image feature. Moreover, this study designs a DFAD to capture important feature information of images and text in two channels. Such a dual-channel backbone improves semantic consistency by utilizing a particular synchronized bi-modal information extraction structure. We have performed experiments on publicly available datasets to prove the effectiveness of our model.
参考文献:
正在载入数据...
