详细信息

基于感知注意力和轻量金字塔融合网络模型的室内场景语义分割方法    

Semantic Segmentation Method of Indoor Scene Based on Perceptual Attention and Lightweight Pyramid Fusion Network Model

文献类型:期刊文献

中文题名:基于感知注意力和轻量金字塔融合网络模型的室内场景语义分割方法

英文题名:Semantic Segmentation Method of Indoor Scene Based on Perceptual Attention and Lightweight Pyramid Fusion Network Model

作者:李钰[1];袁晴龙[1];徐少铭[1];和嘉鹏[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2023

卷号:49

期号:1

起止页码:116

中文期刊名:华东理工大学学报(自然科学版)

外文期刊名:Journal of East China University of Science and Technology

收录:Scopus;北大核心:【北大核心2020】;CSCD:【CSCD_E2023_2024】;

语种:中文

中文关键词:生物实验室场景;感知注意力;轻量金字塔;多尺度特征;语义分割;融合

外文关键词:biological laboratory scene;perceptual attention;lightweight pyramid;multi-scale features;semantic segmentation;fusion

摘要:针对实验室场景理解时存在背景复杂、光照多变等问题,利用RGB信息与深度信息在场景理解中具有互补性的特点,提出了一种感知注意力和轻量空间金字塔融合的网络模型(Perception Attention and Lightweight Spatial Fusion Network,PLFNet)。在该模型的感知注意力模块中,利用RGB图像与深度图像在网络中的权重不同,以加权的方式实现深度信息对RGB信息的多级辅助;在轻量空间金字塔池化模块中,通过增加级联的空洞空间卷积,不但有效地聚集了多尺度特征,而且比传统空间金字塔池化模块的参数量减少了约92%,使RGB信息和深度信息的融合更充分。在两个室内场景公开数据集上的实验结果表明,该模型的表现均优于经典算法。消融实验结果表明,本文模型添加感知注意力模块和轻量空间金字塔池化模块后,平均交并比分别提高了4.3%和3.5%。最后,利用场景较复杂的生物实验室数据集进行测试,结果表明本文模型可以有效地实现对生物实验室的场景理解。
Aiming at the problems of complex background and variable lighting in laboratory scene understanding,this paper proposes a perceptual attention and lightweight spatial fusion network model, PLFNet, by using the complementary characteristics of RGB image information and depth image information in scene understanding. In the perceptual attention module of this model, RGB image and the depth image in the network are used via weighting to implement the multi-level assistance of depth information to the RGB information. In the lightweight spatial pyramid pooling module, by adding cascaded hole spatial convolution, not only the multi-scale features are effectively gathered,but also the parameters are reduced by about 92% compared with the traditional spatial pyramid pooling module, which enables the fusion of RGB information and depth information to fuse more adequately. It is shown via the experiments on two public datasets of indoor scenes that the proposed model performs better than some classic algorithms. The results of ablation experiments verify that the average intersection and union ratio of this model is increased by 4.3%and 3.5%,respectively. Finally, a test based on the dataset of biological laboratory on the more complex scenes is carried out, whose results show that the model can effectively realize the scene understanding of biological laboratory.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心