详细信息
IAM-Edit: Localized Image Editing via Instruction Attention Maps ( SCI-EXPANDED收录)
文献类型:期刊文献
英文题名:IAM-Edit: Localized Image Editing via Instruction Attention Maps
作者:Mao, Shucheng[1];Cheng, Hua[1];Qian, Zehong[1];Ding, Yingying[1];Luo, Shibo[1];Fang, Yiquan[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, 130 Meilong Rd, Shanghai 200237, Peoples R China
年份:2026
外文期刊名:EUROPEAN JOURNAL ON ARTIFICIAL INTELLIGENCE
收录:;WOS:【SCI-EXPANDED(收录号:WOS:001796128300001)】;
基金:The authors received no financial support for the research, authorship, and/or publication of this article.
语种:英文
外文关键词:diffusion models; localized image editing; attention maps
摘要:Diffusion models have demonstrated impressive performance in text-to-image generation and image editing. However, in instruction-based image editing, they often encounter two challenges: (1) inaccurate localization of the editing targets and (2) unintended modifications in nontarget regions. These issues stem from the global processing of diffusion models due to attention mechanisms. To address these limitations, we conduct a systematic analysis of attention maps under editing instructions and design localization instructions to obtain the desired attention. We propose Instruction Attention Maps (IAM)-Edit, a localized image editing framework that explicitly decouples an editing pipeline into two stages: region localization followed by region-aware editing. Specifically, to localize the editing region, a mask is generated by clustering patches of self-attention maps and combining them with the focal points of cross-attention maps under the editing instruction. To preserve nonediting regions, we apply an attention modulation method that adjusts cross-attention weights at each denoising step based on the generated mask, enabling the denoising process to focus on the editing region. Experiments show that IAM-Edit outperforms state-of-the-art methods both qualitatively and quantitatively.
参考文献:
正在载入数据...
