详细信息

ConvPose: A modern pure ConvNet for human pose estimation  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:ConvPose: A modern pure ConvNet for human pose estimation

作者:Niu, Yue[1];Wang, Annan[1];Wang, Xuewu[1,2];Wu, Shengxi[1,2]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, 130 North Meilong Rd, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China

年份:2023

卷号:544

外文期刊名:NEUROCOMPUTING

收录:;EI(收录号:20232014092329);WOS:【SCI-EXPANDED(收录号:WOS:001009045100001)】;

基金:This work was supported in part by the National Natural Science Foundation of China under Grant 62076095, and the authors would like to appreciate the Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology, China.

语种:英文

外文关键词:Deep learning; Convolutional neural network; Transformer; Human pose estimation

摘要:Transformer-based networks almost thoroughly outperformed those based on convolutional neural net-work (ConvNet) and predominate in the field of pose estimation. To get off the hook and resuscitate ConvNets, we propose ConvPose, which is a pure ConvNet that does not utilize conventional improve-ment strategies like attention mechanisms and lightweight approaches, but instead pioneeringly mod-ernizes network structures. The modernization process includes: deepening the stem cell and transition layers, using a separate pointwise convolution layer, adopting a batch normalization (BN) layer after resizing the feature maps, employing large-kernel depthwise separable convolutions and designing re-parameterized-style structures, constructing two consecutive modules that contain a mixer and an inverted bottleneck, etc. All of these designs are similar to the corresponding Transformer architectures, which means translating Transformer-specific components into convolutional variations and incorporat-ing them into a ConvNet. A modern ConvNet not only maintains the simplicity of convolutional, but also takes advantage of Transformers. The experiments show that ConvPose-BL achieves a 76.0 Average Precision (AP) score on the COCO val2017 dataset. ConvPose performs on par or better than the existing representative networks those based on Transformer and ConvNet, and represents slight superiority in terms of speed. (c) 2023 Elsevier B.V. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心