详细信息
Mind-bridge: reconstructing visual images based on diffusion model from human brain activity ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Mind-bridge: reconstructing visual images based on diffusion model from human brain activity
作者:Liu, Qing[1];Zhu, Hongqing[1];Chen, Ning[1];Huang, Bingcang[2];Lu, Weiping[2];Wang, Ying[3]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Gongli Hosp Shanghai Pudong New Area, Dept Radiol, Shanghai 200135, Peoples R China;[3]Gongli Hosp Shanghai Pudong New Area, Shanghai Hlth Commiss, Sino French Cooperat Cent Lab, Key Lab Artificial Intelligence AI Based Managemen, Shanghai 200135, Peoples R China
年份:2024
卷号:18
期号:SUPPL 1
起止页码:953
外文期刊名:SIGNAL IMAGE AND VIDEO PROCESSING
收录:;EI(收录号:20241916033370);WOS:【SCI-EXPANDED(收录号:WOS:001214682700001)】;
基金:This work was supported by the National Nature Science Foundation of China under Grants 61872143, 82372029, 61771196. Discipline Construction of Pudong New Area Health Commission (PWGw2020-01, PWZxk2022-03). Joint Research Project of Pudong New Area Health and Family Planning Commission (PW2021D-14).
语种:英文
外文关键词:Image reconstruction; Diffusion model; Functional magnetic resonance imaging (fMRI); Visual decoding; Variational autoencoder
摘要:Human brain vision is mysterious and complex, and it interprets the world through the connection between the brain and the eyes. In recent years, several methods have relied on fMRI to successfully reconstruct visual images from human brain activity. However, these reconstruction methods focus more on the semantics of the reconstruction image and lack attention to the image structure and foreground targets. To alleviate this problem, we propose a diffusion model-based image reconstruction architecture (Mind-Bridge) that utilizes fMRI to reconstruct visual images from human brain activity. Specifically, we first develop a novel Depth Structure Variational AutoEncoder (DSVAE) to capture image structural information at the initial stage. To obtain more foreground target information, we further introduce Edge estimation through the edge detection operator. In addition, we utilize Contrastive Language Image Pre-training (CLIP) text and image encoders as image and text prompt conditions for visual reconstruction. Finally, our proposed Mind-Bridge utilizes the Versatile Diffusion (VD) to fuse different stages of image information for visual images reconstruction. Qualitative and quantitative analysis results on the challenging Natural Scene Dataset (NSD) show that our proposed Mind-Bridge is effective.
参考文献:
正在载入数据...
