详细信息

Multi-modal self-supervised contrastive representation learning for three-dimensional point cloud understanding  ( EI收录)  

文献类型:期刊文献

英文题名:Multi-modal self-supervised contrastive representation learning for three-dimensional point cloud understanding

作者:Ding, Weichao[1]; Yang, Zehao[1]; Luo, Fei[1]; Gu, Chunhua[1]; Dong, Wenbo[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, 130 Meilong Road, Shanghai, 200237, China

年份:2025

卷号:160

外文期刊名:Engineering Applications of Artificial Intelligence

收录:EI(收录号:20253419024188)

语种:英文

外文关键词:Contrastive Learning - Learning algorithms - Learning systems - Modal analysis - Self-supervised learning - Semantic Web - Supervised learning

摘要:Three-dimensional point cloud understanding is widely used in autonomous driving and has become a hot research topic. However, labeling large scale point cloud data is challenging, and single modal unlabeled data often provides limited information, which severely limits its development. Multi-modal self-supervised learning can learn relevant features from unlabeled data and incorporate more modal information to provide rich semantic information for point cloud data, which shows great potential but is difficult to effectively correct the semantic bias between modalities. In order to address above issues, we propose a multi-modal self-supervised contrastive learning model (named MSCLM) for three-dimensional point cloud understanding, which aims to enhance the understanding of three-dimensional point cloud by jointly learning from multiple modalities of unlabeled data. MSCLM trains on a large amount of point cloud–image–text triplet data and learns better feature representations for point cloud in a self-supervised manner from positive and negative samples. To compensate for semantic biases caused by multiple modalities with different granularity, we design a global point cloud feature fusion mechanism, which reduces the negative impact of biases from pre-trained networks on each modality by integrating global point cloud features with other modal features. To fully integrate the learning information from different modalities, we adopt intra-modal enhanced contrastive learning and cross-modal joint contrastive learning to further strengthen multi-modal information sharing. The experimental results demonstrate the effectiveness of the proposed algorithm and promote the development of autonomous driving. ? 2025 Elsevier Ltd

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心