详细信息
MMNet: Multi-modal multi-stage network for RGB-T image semantic segmentation ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:MMNet: Multi-modal multi-stage network for RGB-T image semantic segmentation
作者:Lan, Xin[1];Gu, Xiaojing[1];Gu, Xingsheng[1]
机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Shanghai 200237, Peoples R China
年份:2022
卷号:52
期号:5
起止页码:5817
外文期刊名:APPLIED INTELLIGENCE
收录:;EI(收录号:20213410815025);WOS:【SCI-EXPANDED(收录号:WOS:000687522700006)】;
基金:This work is supported by National Natural Science Foundation of China under Grant No. 61973122 and 61973120.
语种:英文
外文关键词:Semantic segmentation; Multi-modal; Multi-stage; Feature fusion
摘要:Semantic segmentation is a fundamental task in computer vision. However, most networks are designed for RGB inputs which quality can degrade gracefully under low-level illumination or bad weather conditions. Recent works have achieved promising results by inputting networks with a RGB image and a corresponding registered thermal image together. However, how and when to fuse the features of RGB modality and thermal modality are still remain challenging. In this paper, we propose a Multi-Modal Multi-Stage Network (MMNet) for RGB-T image semantic segmentation. MMNet consists of two stages. Stage 1 extracts features of different modalities separately to avoid cross-modal feature conflicts. Stage 2 fuses representations from the first stage and gradually refines the details. Specifically, Stage 1 has two encoder-decoder sub-networks while Stage 2 has one. As semantic gap exists across encoders and decoders, we propose an Efficient Feature Enhancement Module (EFEM) to bridge the encoder with decoder. Moreover, we deploy a light-weight Mini Refinement Block (MRB) as the encoder at Stage 2 to do the fusion and refinement efficiently. The experimental results demonstrate that our network achieves improved performance while simultaneously being efficient in terms of parameters and FLOPs.
参考文献:
正在载入数据...
