详细信息
PointTransformer: Encoding Human Local Features for Small Target Detection ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:PointTransformer: Encoding Human Local Features for Small Target Detection
作者:Tang, Yudi[1];Wang, Bing[1];He, Wangli[1];Qian, Feng[1];Liu, Zhen[2]
机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, 130 Meilong Rd, Shanghai 200237, Peoples R China;[2]Sinopec Shanghai Petrochem Co Ltd, 48 Jinyi Rd, Shanghai 201512, Peoples R China
年份:2022
卷号:2022
外文期刊名:COMPUTATIONAL INTELLIGENCE AND NEUROSCIENCE
收录:;EI(收录号:20223612691743);WOS:【SCI-EXPANDED(收录号:WOS:000848377300027)】;
基金:AcknowledgmentsThis work was supported by National Key Research and Development Program of China under Grant 2018AAA0101602 and National Natural Science Foundation of China (61922030).
语种:英文
外文关键词:Chemical detection - Chemical plants - Chemicals - Encoding (symbols) - Feature extraction - Object recognition - Signal encoding
摘要:The improvement of small target detection and obscuration handling is the key problem to be solved in the object detection task. In the field operation of chemical plant, due to the occlusion of construction workers and the long distance of surveillance shooting, it often leads to the phenomenon of missed detection. Most of the existing work uses multiple feature fusion strategies to extract different levels of features and then aggregate them into global features, which does not utilize local features and makes it difficult to improve the performance of small target detection. To address this issue, this paper introduces Point Transformer, a transformer encoder, as the core backbone of the object detection framework that first uses a priori information of human skeletal points to obtain local features and then uses both self-attention and cross-attention mechanisms to reconstruct the local features corresponding to each key point. In addition, since the target to be detected is highly correlated with the position of human skeletal points, to further boost Point Transformer's performance, a learnable positional encoding method is proposed by us to highlight the position characteristics of each skeletal point. The proposed model is evaluated on the dataset of field operation in a chemical plant. The results are significantly better than the classical algorithms. It also outperforms state-of-the-art by 12 percent of map points in the small target detection task.
参考文献:
正在载入数据...
