详细信息
MESTrans: Multi-scale embedding spatial transformer for medical image segmentation ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:MESTrans: Multi-scale embedding spatial transformer for medical image segmentation
作者:Liu, Yatong[1];Zhu, Yu[1,2];Xin, Ying[3];Zhang, Yanan[3];Yang, Dawei[2,4];Xu, Tao[3]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Shanghai Engn Res Ctr Internet Things Resp Med, Shanghai 200237, Peoples R China;[3]Qingdao Univ, Dept Pulm & Crit Care Med, Affiliated Hosp, Qingdao 266000, Shandong, Peoples R China;[4]Fudan Univ, Zhongshan Hosp, Dept Pulm & Crit Care Med, Shanghai 200032, Peoples R China
年份:2023
卷号:233
外文期刊名:COMPUTER METHODS AND PROGRAMS IN BIOMEDICINE
收录:;EI(收录号:20231413853143);WOS:【SCI-EXPANDED(收录号:WOS:000961580100001)】;
基金:This work was supported in part by the National Natural Science Foundation of China (Grant No. 82170110 ), the Science and Technology Commission of Shanghai Municipality (Grant Nos. 20DZ22544000, 20DZ2261200, ZD2021CY001), and the Fujian Province Department of Science and Technology (Grant No. 2022D014 )
语种:英文
外文关键词:Computer-aided diagnosis; COVID-19; Medical image segmentation; Transformer
摘要:Background and objective: Transformers profiting from global information modeling derived from the self -attention mechanism have recently achieved remarkable performance in computer vision. In this study, a novel transformer-based medical image segmentation network called the multi-scale embedding spatial transformer (MESTrans) was proposed for medical image segmentation. Methods: First, a dataset called COVID-DS36 was created from 4369 computed tomography (CT) im-ages of 36 patients from a partner hospital, of which 18 had COVID-19 and 18 did not. Subsequently, a novel medical image segmentation network was proposed, which introduced a self-attention mechanism to improve the inherent limitation of convolutional neural networks (CNNs) and was capable of adap-tively extracting discriminative information in both global and local content. Specifically, based on U-Net, a multi-scale embedding block (MEB) and multi-layer spatial attention transformer (SATrans) structure were designed, which can dynamically adjust the receptive field in accordance with the input content. The spatial relationship between multi-level and multi-scale image patches was modeled, and the global context information was captured effectively. To make the network concentrate on the salient feature region, a feature fusion module (FFM) was established, which performed global learning and soft selec-tion between shallow and deep features, adaptively combining the encoder and decoder features. Four datasets comprising CT images, magnetic resonance (MR) images, and H&E-stained slide images were used to assess the performance of the proposed network. Results: Experiments were performed using four different types of medical image datasets. For the COVID-DS36 dataset, our method achieved a Dice similarity coefficient (DSC) of 81.23%. For the GlaS dataset, 89.95% DSC and 82.39% intersection over union (IoU) were obtained. On the Synapse dataset, the average DSC was 77.48% and the average Hausdorff distance (HD) was 31.69 mm. For the I2CVB dataset, 92.3% DSC and 85.8% IoU were obtained. Conclusions: The experimental results demonstrate that the proposed model has an excellent generaliza-tion ability and outperforms other state-of-the-art methods. It is expected to be a potent tool to assist clinicians in auxiliary diagnosis and to promote the development of medical intelligence technology. (c) 2023 Elsevier B.V. All rights reserved.
参考文献:
正在载入数据...
