详细信息

Semantic Mask Reconstruction and Category Semantic Learning for few-shot image generation  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Semantic Mask Reconstruction and Category Semantic Learning for few-shot image generation

作者:Xiao, Ting[1,2];Cai, Yunjie[2];Guan, Jiaoyan[2];Wang, Zhe[1,2]

机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2025

卷号:183

外文期刊名:NEURAL NETWORKS

收录:;EI(收录号:20245017498779);WOS:【SCI-EXPANDED(收录号:WOS:001374259900001)】;

基金:Acknowledgments This work is supported by the National Natural Science Foundation of China under Grant No. 62476087 and No. 62306115, and the Na-tional Key Research and Development Program of China (2022YFB3203500) .

语种:英文

外文关键词:Few-shot image generation; Generative adversarial network; Mask reconstruct; Semantic learning

摘要:Few-shot image generation aims at generating novel images for the unseen category when given K images from the same category. Despite significant advancements in existing few-shot image generation methods, great challenges remain regarding the quality and diversity of the generated images. This issue stems from the model's struggle to fully comprehend the semantic content of images and extract sufficiently semantic representations. To address these issues, we propose a semantic mask reconstruction (SMR) and category semantic learning (CSL) method for few-shot image generation. Specifically, SMR performs mask reconstruction in a high-level semantic space and designs a strategy for dynamically adjusting the mask ratio, which increases the difficulty of the generation tasks by gradually increasing the mask ratio to enhance the learning ability of the discriminator, thereby prompting the generator to learn more critical features relevant to the generation task. In addition, CSL introduces a triplet loss to optimize the distance between the generated image, its corresponding input image, and input images of other categories. This encourages the generative model to discern subtle differences between categories, thereby achieving more fine-grained generation and improving the fidelity of generated images. Both SMR and CSL can function as plug-and-play modules. Extensive experimental results across three standard datasets demonstrate that the SMR-CSL outperforms other methods in terms of the quality and diversity of the generated images. Furthermore, the results of downstream classification experiments verify that the images generated by the proposed method can effectively assist downstream classification tasks.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心