详细信息

CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting Mechanism  ( CPCI-S收录)  

文献类型:会议论文

英文题名:CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting Mechanism

作者:Zhu, Ziming[1];Zhu, Yu[1];Chen, Jiahao[1];Ling, Xiaofeng[1,2];Chen, Huanlei[3];Sun, Lihua[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Shanghai Key Lab Intelligent Sensing & Detect Tec, Shanghai, Peoples R China;[3]Shanghai Motor Vehicle Inspect Certificat & Tech, Shanghai, Peoples R China

会议论文集:42nd International Conference on Machine Learning-ICML-Annual

会议日期:JUL 13-19, 2025

会议地点:Vancouver, CANADA

语种:英文

摘要:Recently, image-based 3D semantic occupancy prediction has become a hot topic in 3D scene understanding for autonomous driving. Compared with the bounding box form of 3D object detection, the ability to describe the fine-grained contours of any obstacles in the scene is the key insight of voxel occupancy representation, which facilitates subsequent tasks of autonomous driving. In this work, we propose CSV-Occ to address the following two challenges: (1) Existing methods fuse temporal information based on the attention mechanism, but are limited by high complexity. We extend the state space model to support multi-input sequence interaction and conduct temporal modeling in a cascaded architecture, thereby reducing the computational complexity from quadratic to linear. (2) Existing methods are limited by semantic ambiguity, resulting in the centers of foreground objects often being predicted as empty voxels. We enable the model to explicitly vote for the instance center to which the voxels belong and spontaneously learn to utilize the other voxel features of the same instance to update the semantics of the internal vacancies of the objects from coarse to fine. Experiments on the Occ3D-nuScenes dataset show that our method achieves state-of-the-art in camera-based 3D semantic occupancy prediction and also performs well on lidar point cloud semantic segmentation on the nuScenes dataset. Code will be available at https://github.com/ZeaZoM/CSV-Occ.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心