详细信息

PointTransformer: Encoding Human Local Features for Small Target Detection  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:PointTransformer: Encoding Human Local Features for Small Target Detection

作者:Tang, Yudi[1];Wang, Bing[1];He, Wangli[1];Qian, Feng[1];Liu, Zhen[2]

机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, 130 Meilong Rd, Shanghai 200237, Peoples R China;[2]Sinopec Shanghai Petrochem Co Ltd, 48 Jinyi Rd, Shanghai 201512, Peoples R China

年份:2022

卷号:2022

外文期刊名:COMPUTATIONAL INTELLIGENCE AND NEUROSCIENCE

收录:;EI(收录号:20223612691743);WOS:【SCI-EXPANDED(收录号:WOS:000848377300027)】;

基金:AcknowledgmentsThis work was supported by National Key Research and Development Program of China under Grant 2018AAA0101602 and National Natural Science Foundation of China (61922030).

语种:英文

外文关键词:Chemical detection - Chemical plants - Chemicals - Encoding (symbols) - Feature extraction - Object recognition - Signal encoding

摘要:The improvement of small target detection and obscuration handling is the key problem to be solved in the object detection task. In the field operation of chemical plant, due to the occlusion of construction workers and the long distance of surveillance shooting, it often leads to the phenomenon of missed detection. Most of the existing work uses multiple feature fusion strategies to extract different levels of features and then aggregate them into global features, which does not utilize local features and makes it difficult to improve the performance of small target detection. To address this issue, this paper introduces Point Transformer, a transformer encoder, as the core backbone of the object detection framework that first uses a priori information of human skeletal points to obtain local features and then uses both self-attention and cross-attention mechanisms to reconstruct the local features corresponding to each key point. In addition, since the target to be detected is highly correlated with the position of human skeletal points, to further boost Point Transformer's performance, a learnable positional encoding method is proposed by us to highlight the position characteristics of each skeletal point. The proposed model is evaluated on the dataset of field operation in a chemical plant. The results are significantly better than the classical algorithms. It also outperforms state-of-the-art by 12 percent of map points in the small target detection task.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心