详细信息

Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation  ( EI收录)  

文献类型:期刊文献

英文题名:Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation

作者:Khalid, Mahzaib[1];Ying, Fangli[1];Atef, Al-Garadi Ahmed Mohammed[1];Phaphuangwittayakul, Aniwat[2];Dhuny, Riyad[3]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci, Shanghai 200237, Peoples R China;[2]Chiang Mai Univ, Int Coll Digital Innovat, Chiang Mai 50200, Thailand;[3]Univ Technol, Dept Creat Arts Film & Media Technol, Port Louis 11134, Mauritius

年份:2026

卷号:12

期号:7

外文期刊名:JOURNAL OF IMAGING

收录:EI(收录号:20263121194571);Scopus(收录号:2-s2.0-105045814871);WOS:【ESCI(收录号:WOS:001832137200001)】;

基金:This research was funded by the Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China (Grant No. JYB2025XDXM402) and the National Major Scientific Instruments and Equipment Development Project of the National Natural Science Foundation of China (Grant No. 32327801).

语种:英文

外文关键词:diffusion models; stable diffusion; concept-vector learning; prompt-guided learning; semantic manipulation; image-to-image editing; bottleneck feature injection; abstract visual concepts

摘要:Diffusion-based generative models achieve high-fidelity image synthesis; however, controlling internal representations for abstract visual concepts remains challenging due to the ambiguity of textual descriptions. In this work, we propose a prompt-guided concept-vector learning framework for the controllable manipulation of such concepts without requiring external human-annotated image pairs, segmentation masks, identity labels, or manually annotated editing targets. The method introduces a learnable concept vector optimized in the bottleneck (mid-block) feature space of a pretrained Stable Diffusion U-Net, while keeping all pretrained model parameters frozen. A multi-prompt data generation strategy based on paired positive and neutral prompts provides weak semantic guidance for capturing the target concept direction and reducing dependence on a single prompt formulation. The learned vector is further applied in an image-to-image setting through controlled noise injection and concept-guided denoising, enabling the semantic modification of real images while preserving structural content. The concept strength is controlled by a scaling parameter gamma , while the image-to-image noise strength is controlled by beta , allowing for a practical balance between semantic modification and structural fidelity. Experiments are conducted on two main abstract concepts, perfect skin and peaceful lake, with additional qualitative analysis on subjective portrait-level concepts. Quantitative evaluation using SSIM, LPIPS, and CLIP similarity demonstrates that the proposed method improves semantic alignment while maintaining structural preservation compared with Stable Diffusion image-to-image baselines. A human preference study further shows that concept-injected outputs are preferred in 76.0% of responses for perfect skin and 85.7% for peaceful lake. Ablation studies further demonstrate the controllability and robustness of the proposed framework. Overall, the method provides a simple and parameter-efficient approach for interpretable concept-level manipulation in diffusion models.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心