详细信息
Rethinking Mobilevitv2 for Real-Time Semantic Segmentation ( EI收录)
文献类型:期刊文献
英文题名:Rethinking Mobilevitv2 for Real-Time Semantic Segmentation
作者:Zhang, Zhengbin[1]; Xu, Zhenhao[1]; Gu, Xingsheng[1]; Xiong, Juan[2]
机构:[1] East China University of Science and Technology, School of Information Science and Engineering, Shanghai, China; [2] University of Shanghai for Science and Technology, School of Optical Electircal and Computer Engineering, Shanghai, China
年份:2023
起止页码:57
外文期刊名:2023 4th International Conference on Computers and Artificial Intelligence Technology, CAIT 2023
收录:EI(收录号:20241615928170)
语种:英文
外文关键词:Computer vision - Economic and social effects - Semantics
摘要:Semantic segmentation empowers various real-world applications. Nevertheless, the substantial computational cost, such as O(k2) time complexity associated with the number of tokens in multi-headed self-attention, poses challenges for deploying these models on edge devices with constrained hardware resources. This paper introduces a novel family of backbones designed for real-time semantic segmentation, referred to as Linear and Re-parameter Vision Transformer (LARFormer). Particularly, we introduce a Re-parameter Mobile Block (RMB), which employs three branches during training and a single branch during inference. Furthermore, we introduce Linear Separable Self-Attention(LSSA), which reduces the computational complexity from O(k2) to O(k). Extensive experiments on the ADE20K dataset and Pascal VOC 2012 dataset demonstrate the effectiveness of the proposed LARFormer by achieving a promising trade-off between segmentation accuracy and inference speed. ? 2023 IEEE.
参考文献:
正在载入数据...
