详细信息
Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation ( EI收录)
文献类型:期刊文献
英文题名:Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation
作者:Khalid, Mahzaib[1];Ying, Fangli[1];Atef, Al-Garadi Ahmed Mohammed[1];Phaphuangwittayakul, Aniwat[2];Dhuny, Riyad[3]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci, Shanghai 200237, Peoples R China;[2]Chiang Mai Univ, Int Coll Digital Innovat, Chiang Mai 50200, Thailand;[3]Univ Technol, Dept Creat Arts Film & Media Technol, Port Louis 11134, Mauritius
年份:2026
卷号:12
期号:7
外文期刊名:JOURNAL OF IMAGING
收录:EI(收录号:20263121194571);Scopus(收录号:2-s2.0-105045814871);WOS:【ESCI(收录号:WOS:001832137200001)】;
基金:This research was funded by the Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China (Grant No. JYB2025XDXM402) and the National Major Scientific Instruments and Equipment Development Project of the National Natural Science Foundation of China (Grant No. 32327801).
语种:英文
外文关键词:diffusion models; stable diffusion; concept-vector learning; prompt-guided learning; semantic manipulation; image-to-image editing; bottleneck feature injection; abstract visual concepts
摘要:Diffusion-based generative models achieve high-fidelity image synthesis; however, controlling internal representations for abstract visual concepts remains challenging due to the ambiguity of textual descriptions. In this work, we propose a prompt-guided concept-vector learning framework for the controllable manipulation of such concepts without requiring external human-annotated image pairs, segmentation masks, identity labels, or manually annotated editing targets. The method introduces a learnable concept vector optimized in the bottleneck (mid-block) feature space of a pretrained Stable Diffusion U-Net, while keeping all pretrained model parameters frozen. A multi-prompt data generation strategy based on paired positive and neutral prompts provides weak semantic guidance for capturing the target concept direction and reducing dependence on a single prompt formulation. The learned vector is further applied in an image-to-image setting through controlled noise injection and concept-guided denoising, enabling the semantic modification of real images while preserving structural content. The concept strength is controlled by a scaling parameter gamma , while the image-to-image noise strength is controlled by beta , allowing for a practical balance between semantic modification and structural fidelity. Experiments are conducted on two main abstract concepts, perfect skin and peaceful lake, with additional qualitative analysis on subjective portrait-level concepts. Quantitative evaluation using SSIM, LPIPS, and CLIP similarity demonstrates that the proposed method improves semantic alignment while maintaining structural preservation compared with Stable Diffusion image-to-image baselines. A human preference study further shows that concept-injected outputs are preferred in 76.0% of responses for perfect skin and 85.7% for peaceful lake. Ablation studies further demonstrate the controllability and robustness of the proposed framework. Overall, the method provides a simple and parameter-efficient approach for interpretable concept-level manipulation in diffusion models.
参考文献:
正在载入数据...
