详细信息

Deep Features Homography Transformation Fusion Network-A Universal Foreground Segmentation Algorithm for PTZ Cameras and a Comparative Study  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Deep Features Homography Transformation Fusion Network-A Universal Foreground Segmentation Algorithm for PTZ Cameras and a Comparative Study

作者:Tao, Ye[1];Ling, Zhihao[1]

机构:[1]East China Univ Sci & Technol, Minist Educ, Key Lab Adv Control & Optimizat Chem Proc, Shanghai 200237, Peoples R China

年份:2020

卷号:20

期号:12

起止页码:1

外文期刊名:SENSORS

收录:;EI(收录号:20202508854441);WOS:【SCI-EXPANDED(收录号:WOS:000554582800001)】;

基金:This research was funded by the Fundamental Research Funds for the Central Universities grant number 222201917006.

语种:英文

外文关键词:moving object segmentation; PTZ camera; convolutional neural network; image alignment

摘要:The foreground segmentation method is a crucial first step for many video analysis methods such as action recognition and object tracking. In the past five years, convolutional neural network based foreground segmentation methods have made a great breakthrough. However, most of them pay more attention to stationary cameras and have constrained performance on the pan-tilt-zoom (PTZ) cameras. In this paper, an end-to-end deep features homography transformation and fusion network based foreground segmentation method (HTFnetSeg) is proposed for surveillance videos recorded by PTZ cameras. In the kernel of HTFnetSeg, there is the combination of an unsupervised semantic attention homography estimation network (SAHnet) for frames alignment and a spatial transformed deep features fusion network (STDFFnet) for segmentation. The semantic attention mask in SAHnet reinforces the network to focus on background alignment by reducing the noise that comes from the foreground. STDFFnet is designed to reuse the deep features extracted during the semantic attention mask generation step by aligning the features rather than only the frames, with a spatial transformation technique in order to reduce the algorithm complexity. Additionally, a conservative strategy is proposed for the motion map based post-processing step to further reduce the false positives that are brought by semantic noise. The experiments on both CDnet2014 and Lasiesta show that our method outperforms many state-of-the-art methods, quantitively and qualitatively.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心