详细信息

Enhancing Conditional Molecular Generation With Pretrained SMILES Transformer and Contrastive Representation Learning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Enhancing Conditional Molecular Generation With Pretrained SMILES Transformer and Contrastive Representation Learning

作者:Hu, Dongcheng[1];Peng, Xin[2];He, Song[1];Wei, Zhangpeng[1];Zhong, Weimin[3];Qian, Feng[2]

机构:[1]East China Univ Sci & Technol, Minist Educ, Key Lab Smart Mfg Energy Chem Proc, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, State Key Lab Ind Control Technol, Shanghai 200237, Peoples R China;[3]East China Univ Sci & Technol, State Key Lab Chem Engn & Low Carbon Technol, Shanghai 200237, Peoples R China

年份:2026

外文期刊名:IEEE TRANSACTIONS ON CYBERNETICS

收录:;EI(收录号:20262621011958);WOS:【SCI-EXPANDED(收录号:WOS:001797739900001)】;

基金:This work was supported in part by the Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China under Grant JYB2025XDXM402, in part by the National Natural Science Foundation of China under Grant 62303186, in part by the Program of Introducing Talents of Discipline to Universities (the 111 Project) under Grant B17017, and in part by the Fundamental Research Funds for the Central Universities.

语种:英文

外文关键词:Modeling; Contrastive learning; Training; Sequences; Sequential analysis; Chemicals; Modules (abstract algebra); Learning (artificial intelligence); Transformers; Educational institutions; Gumbel-Softmax sampling; molecular generation; simplified molecular input line entry system (SMILES) Transformer

摘要:Predicting molecular structures with target properties and specific reaction conditions is a critical task in drug discovery and material science. Simplified molecular input line entry system (SMILES) representation, widely used in molecular generation tasks, encodes molecular graph structures as linear sequences of characters. This redundancy introduces ambiguity in feature mapping, potentially confusing generation models. To address this issue, we propose a pretrained contrastive learning-based SMILES Transformer Encoder (Contras-STE) to capture invariant, high-level features. This approach guides the SMILES generation model in learning the inherent structure of SMILES while maintaining a differentiable conversion process. To avoid nondifferentiability, we sample the generated SMILES from the token distribution using Gumbel-Softmax and input them into Contras-STE to compute the InfoNCE loss, which quantifies the similarity between the generated SMILES, target SMILES, and irrelevant SMILES. We test Contras-STE on MoleculeNet benchmarks against existing fingerprint-based methods and RNN-based methods, and the results show that Contras-STE outperforms other methods in most cases. To evaluate the performance of the SMILES generation model with Contras-STE, we test models on the OSDAs prediction task under synthesis conditions and the properties of zeolites production, and the results demonstrate that Contras-STE can highly improve the validity rate and novelty rate of generated SMILES.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心