详细信息

Rethinking Mobilevitv2 for Real-Time Semantic Segmentation  ( EI收录)  

文献类型:期刊文献

英文题名:Rethinking Mobilevitv2 for Real-Time Semantic Segmentation

作者:Zhang, Zhengbin[1]; Xu, Zhenhao[1]; Gu, Xingsheng[1]; Xiong, Juan[2]

机构:[1] East China University of Science and Technology, School of Information Science and Engineering, Shanghai, China; [2] University of Shanghai for Science and Technology, School of Optical Electircal and Computer Engineering, Shanghai, China

年份:2023

起止页码:57

外文期刊名:2023 4th International Conference on Computers and Artificial Intelligence Technology, CAIT 2023

收录:EI(收录号:20241615928170)

语种:英文

外文关键词:Computer vision - Economic and social effects - Semantics

摘要:Semantic segmentation empowers various real-world applications. Nevertheless, the substantial computational cost, such as O(k2) time complexity associated with the number of tokens in multi-headed self-attention, poses challenges for deploying these models on edge devices with constrained hardware resources. This paper introduces a novel family of backbones designed for real-time semantic segmentation, referred to as Linear and Re-parameter Vision Transformer (LARFormer). Particularly, we introduce a Re-parameter Mobile Block (RMB), which employs three branches during training and a single branch during inference. Furthermore, we introduce Linear Separable Self-Attention(LSSA), which reduces the computational complexity from O(k2) to O(k). Extensive experiments on the ADE20K dataset and Pascal VOC 2012 dataset demonstrate the effectiveness of the proposed LARFormer by achieving a promising trade-off between segmentation accuracy and inference speed. ? 2023 IEEE.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心