详细信息

OnionNet: Single-View Depth Prediction and Camera Pose Estimation for Unlabeled Video  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:OnionNet: Single-View Depth Prediction and Camera Pose Estimation for Unlabeled Video

作者:Gu, Tianhao[1,2];Wang, Zhe[1,2];Li, Dongdong[2];Yang, Hai[2];Du, Wenli[1];Zhou, Yangming[2]

机构:[1]East China Univ Sci & Technol, Key Lab Adv Control & Optimizat Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2021

卷号:13

期号:4

起止页码:995

外文期刊名:IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS

收录:;EI(收录号:20205209676664);WOS:【SCI-EXPANDED(收录号:WOS:000728925200026)】;

基金:This work was supported in part by the Shanghai Science and Technology Program "Distributed and Generative Few-Shot Algorithm and Theory Research" under Grant 20511100600; in part by the Natural Science Foundation of China under Grant 62076094 and Grant 61806078; in part by the National Major Scientific and Technological Special Project for "Significant New Drugs Development" under Grant 2019ZX09201004; in part by the Zhejiang Lab under Grant 2019ND0AB01; in part by the National Science Foundation of China for Distinguished Young Scholars under Grant 61725301; and in part by the National Key Research and Development Project of Ministry of Science and Technology of China under Grant 2018AAA0101302.

语种:英文

外文关键词:Cameras; Training; Pose estimation; Geometry; Robustness; Task analysis; Decoding; Camera pose estimation; multitask learning; single-view depth prediction; unsupervised learning

摘要:In real scenes, humans can easily infer their positions and distances from other objects with their own eyes. To make the robots have the same visual ability, this article presents an unsupervised OnionNet framework, including LeafNet and ParachuteNet, for single-view depth prediction and camera pose estimation. In OnionNet, for speeding up OnionNet's convergence and concretizing objects against the gradient locality and moving objects in videos, LeafNet adopts two decoders and enhanced upconvolution modules. Meanwhile, for improving the robustness of fast camera movement and rotation, ParachuteNet uses and integrates three pose networks to estimate multiview camera pose parameters by combining with the modified image preprocess. Different from existing methods, single-view depth prediction and camera pose estimation are trained view by view, where the variations between views is gradual reduction of view range and outer pixels disappear in next view, similar to onion peeling. Moreover, the LeafNet is optimized with pose parameter from each pose network in turn. Experimental results on the KITTI data set show the outstanding effectiveness of our method: single-view depth performs better than most supervised and unsupervised methods which contain two same subtasks, and pose estimation gets the state-of-the-art performance compared with existing methods under the comparable input settings.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心