详细信息

FedVOD: A two-stage video object detector training framework based on federated unsupervised learning and feature post-processing  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:FedVOD: A two-stage video object detector training framework based on federated unsupervised learning and feature post-processing

作者:Hu, Han[1,2];Du, Wenli[1,2,3,4];Wang, Bing[1,2];Qian, Feng[1,2,3,4]

机构:[1]East China Univ Sci & Technol, State Key Lab Ind Control Technol, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[3]Huzhou Inst Ind Control Technol, Huzhou 313099, Peoples R China;[4]East China Univ Sci & Technol, Engn Res Ctr Proc Syst Engn, Minist Educ, Shanghai 200237, Peoples R China

年份:2025

卷号:315

外文期刊名:KNOWLEDGE-BASED SYSTEMS

收录:;EI(收录号:20251018000508);WOS:【SCI-EXPANDED(收录号:WOS:001442059900001)】;

基金:This work was supported by the National Key Research and Development Program of China (2022YFB3305900) , National Natural Science Foundation of China (62293504) , Major Science and Technology Projects of Longmen Laboratory (No. LMZDXM202206) , the Programme of Introducing Talents of Discipline to Universities (the 111 Project) under Grant B17017 and Fundamental Research Funds for the Central Universities.

语种:英文

外文关键词:Video object detection; Federated unsupervised learning; Adaptive keyframe scheduling; Feature post-processing

摘要:Most existing video object detection (VOD) methods are developed based on centrally stored and densely labeled video data, which is not always practical. In many real-world scenarios (e.g., video surveillance and autonomous driving), video data is decentralized and lacks annotation. Centralizing this data not only increases communication and storage costs but also poses risks of privacy breaches. To address the above issues, this paper proposes a novel two-stage video object detector training framework called FedVOD. In this framework, we first introduce federated unsupervised learning to train feature extractors with multi-client unannotated data while preserving privacy. To overcome the heterogeneity and imbalance challenges of decentralized data, we develop a probability-guided communication protocol and DE-Avg model aggregation algorithm for federated system. Subsequently, an adaptive keyframe scheduler is devised to flexibly adjust the keyframe interval, thereby enhancing the rationality of keyframe selection. On this basis, to effectively strike the balance between the speed and accuracy of VOD, we perform different levels of feature extraction for keyframes and non-keyframes, and design two distinct feature post-processing workflows for the extracted features. For different data distributions, experimental results on the ImageNet VID dataset demonstrate the effectiveness and superiority of FedVOD. In addition, the detection performance in a security surveillance system reflects the potential of the framework to be applied to complex real-world scenarios.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心