详细信息
CJE-TIG: Zero-shot cross-lingual text-to-image generation by Corpora-based Joint Encoding ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:CJE-TIG: Zero-shot cross-lingual text-to-image generation by Corpora-based Joint Encoding
作者:Zhang, Han[1];Yang, Suyi[2];Zhu, Hongqing[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Kings Coll London, Dept Math Nat Math & Engn Sci, London WC2R 2LS, England
年份:2022
卷号:239
外文期刊名:KNOWLEDGE-BASED SYSTEMS
收录:;EI(收录号:20220111430605);WOS:【SCI-EXPANDED(收录号:WOS:000788633300003)】;
基金:Acknowledgments This work was supported by the National Natural Science Foundation of China under Grant 61872143.
语种:英文
外文关键词:Cross-lingual pre-training; Text-to-image synthesis; Universal contextual word vector space; Semantic alignment; Joint adversarial training
摘要:Recently, text-to-Image (T2I) generation has been well developed by improving synthesis authenticity, text-consistency and generation diversity. However, large amount of pairwise image-text data required restricts generalization of synthesis models only to its pre-trained language. In this paper, a cross-lingual pre-training method is proposed to adapt target low-resource language to pre-trained generative models. As far as we known, this is the first time that arbitrary input languages could access T2I generation. This joint encoding scheme fulfills both universal and visual semantic alignment. With any prepared GAN-based T2I framework, pre-trained source encoder model could be easily fine-tuned to construct target encoder model and hence entirely enable transfer of T2I synthesis ability between languages. After that, a semantic-level alignment independent of source T2I structure is established to guarantee optimal text consistency and detail generation. Different from monolingual T2I methods that apply discriminator to enhance generation quality, we use an adversarial training scheme that optimizes the sentence-level alignment along with the word-level alignment with a self-attention mechanism. Considering of training for low-resource languages lack of parallel texts in practice, target input embedding is designed available for zero-shot learning. Experimental results prove robustness of the proposed cross-lingual T2I pre-training on multiple downstream generative models and target languages applied.(c) 2021 Elsevier B.V. All rights reserved.
参考文献:
正在载入数据...
