详细信息

基于多阶门控聚合网络的光学化学结构识别    

Optical Chemical Structure Recognition Based on Multi-order Gated Aggregation Network

文献类型:期刊文献

中文题名:基于多阶门控聚合网络的光学化学结构识别

英文题名:Optical Chemical Structure Recognition Based on Multi-order Gated Aggregation Network

作者:林帆[1];李建华[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2025

卷号:51

期号:8

起止页码:364

中文期刊名:计算机工程

外文期刊名:Computer Engineering

收录:;北大核心:【北大核心2023】;

基金:国家科技重大专项(2018ZX09735002)。

语种:中文

中文关键词:光学化学结构识别;编码解码架构;深度学习;SMILES表达式;多阶门控聚合网络

外文关键词:Optical Chemical Structure Recognition(OCSR);encoder-decoder architecture;deep learning;SMILES notation;Multi-order gated aggregation Network(MogaNet)

摘要:在光学化学结构识别(OCSR)领域,现有基于深度学习的模型通常依赖于卷积神经网络(CNN)或视觉Transformer进行视觉特征提取,并采用Transformer进行序列解码。这些模型虽然有效,但仍受限于图像特征提取能力和解码时位置编码的精确性,从而影响识别效率。针对这些限制,将多阶门控聚合网络(MogaNet)和引入相对位置编码的Transformer构成的编码解码架构用于OCSR领域,提出一种基于多阶门控聚合网络的光学化学结构识别模型。该模型首先在图像特征提取时通过MogaNet空间聚合模块,捕获多尺度特征并减少特征冗余,并且通过MogaNet通道聚合模块改善通道维度的多样性;其次在序列解码时采用引入相对位置编码的Transformer作为解码器,精准捕捉序列单词之间的相对位置关系。为了训练和验证该模型,构建一个包含40万个分子的化学结构数据集,其中包含Markush结构与非Markush结构。实验结果表明,该模型的准确率达到了92.36%,优于其他现有的模型。
In the field of Optical Chemical Structure Recognition(OCSR),current deep-learning-based models predominantly utilize Convolutional Neural Networks(CNNs)or Vision Transformers for visual feature extraction and Transformers for sequence decoding.Although these models are effective,they are still limited by their ability to extract image features and the accuracy of position encoding during decoding,which affect the recognition efficiency.In response to these limitations,this study uses an encoder-decoder architecture composed of a Multiorder gated aggregation Network(MogaNet)and a Transformer,which introduces relative positional encoding in the OCSR field,and proposes an optical chemical structure recognition model based on MogaNet.First,the model captures multiscale features,reduces feature redundancy using the MogaNet spatial aggregation module during image feature extraction,and improves channel dimension diversity using the MogaNet channel aggregation module.Second,during sequence decoding,a Transformer with relative positional encoding is used as the decoder to accurately capture the relative positional relationships between words.To train and validate this model,a chemical structure dataset containing 400000 molecular structures is constructed,which includes both Markush and non-Markush structures.Experimental results demonstrate that the model achieves an accuracy of 92.36%,outperforming other models.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心