详细信息
GCLmf: A Novel Molecular Graph Contrastive Learning Framework Based on Hard Negatives and Application in Toxicity Prediction ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:GCLmf: A Novel Molecular Graph Contrastive Learning Framework Based on Hard Negatives and Application in Toxicity Prediction
作者:Yu, Xinxin[1];Chen, Yuanting[1];Chen, Long[1];Li, Weihua[1];Wang, Yuhao[1];Tang, Yun[1];Liu, Guixia[1]
机构:[1]East China Univ Sci & Technol, Shanghai Frontiers Sci Ctr Optogenet Tech Cell Met, Sch Pharm, Shanghai Key Lab New Drug Design, 130 Meilong Rd, Shanghai 200237, Peoples R China
年份:2025
卷号:44
期号:1
外文期刊名:MOLECULAR INFORMATICS
收录:;EI(收录号:20244217226066);WOS:【SCI-EXPANDED(收录号:WOS:001334579200001)】;
基金:This work was supported by the National Key Research and Development Program of China (Grant 2019YFA0904800), the National Natural Science Foundation of China (Grants 82173746 and 82273858) and Shanghai Frontiers Science Center of Optogenetic Techniques for Cell Metabolism (Shanghai Municipal Education Commission, Grant 2021 Sci & Tech 03-28).
语种:英文
外文关键词:chemical toxicity; graph contrastive learning; GNN; hard negatives; molecular representation; self-supervised learning
摘要:In silico methods for prediction of chemical toxicity can decrease the cost and increase the efficiency in the early stage of drug discovery. However, due to low accessibility of sufficient and reliable toxicity data, constructing robust and accurate prediction models is challenging. Contrastive learning, a type of self-supervised learning, leverages large unlabeled data to obtain more expressive molecular representations, which can boost the prediction performance on downstream tasks. While molecular graph contrastive learning has gathered growing attentions, current models neglect the quality of negative data set. Here, we proposed a self-supervised pretraining deep learning framework named GCLmf. We first utilized molecular fragments that meet specific conditions as hard negative samples to boost the quality of the negative set and thus increase the difficulty of the proxy tasks during pre-training to learn informative representations. GCLmf has shown excellent predictive power on various molecular property benchmarks and demonstrates high performance in 33 toxicity tasks in comparison with multiple baselines. In addition, we further investigated the necessity of introducing hard negatives in model building and the impact of the proportion of hard negatives on the model. In this paper, we proposed a self-supervised pretraining deep learning framework named GCLmf. GCLmf has shown excellent predictive power on various molecular property benchmarks and demonstrates high performance in 36 toxicity tasks in comparison with multiple baselines. image
参考文献:
正在载入数据...
