详细信息
以离子液体密度为例的分子性质预测模型建模方法探讨 ( EI收录)
A critical discussion on developing molecular property prediction models:density of ionic liquids as example
文献类型:期刊文献
中文题名:以离子液体密度为例的分子性质预测模型建模方法探讨
英文题名:A critical discussion on developing molecular property prediction models:density of ionic liquids as example
作者:陈家辉[1];杨鑫泽[1];陈顾中[1];宋震[1];漆志文[1]
机构:[1]华东理工大学化工学院,化学工程联合国家重点实验室,上海200237
年份:2023
卷号:74
期号:2
起止页码:630
中文期刊名:化工学报
外文期刊名:CIESC Journal
收录:CSTPCD;;EI(收录号:20231413833312);Scopus;北大核心:【北大核心2020】;CSCD:【CSCD2023_2024】;
语种:中文
中文关键词:分子性质预测;模型;数据集划分;交叉验证;算法;离子液体;密度
外文关键词:molecular property prediction;modelling;dataset partitioning;cross-validation;algorithms;ionic liquids;density
摘要:分子性质预测模型是针对特定应用需求筛选设计化学品的有力工具,然而诸多相关建模过程中的测试集划分、交叉验证、算法选择等关键环节普遍存在严谨性不足的问题,模型真实预测性能难以保证。以基团贡献法预测离子液体密度为例,探讨了分子性质预测模型建模过程中数据集划分和交叉验证的重要性,提出了自动基团划分方法并研究了数据集中基团涉及分子个数对预测精度的影响。通过对比五种回归算法(多重线性回归、岭回归、随机森林、支持向量机、神经网络),基于岭回归的基团贡献模型预测性能最佳,在由1078种离子液体、共计23034个数据点组成的数据集上得到的平均相对误差为1.88%。
Molecular property prediction models are powerful tools for screening or designing chemicals to meet specific application requirements.However,many key aspects in model development such as the size and diversity of dataset,test set partitioning method,cross-validation,and algorithm selection are not treated with enough rigor,which could lead to doubtful estimation of the true predictive performance of models.Taking the group contribution method to predict the density of ionic liquids as an example,the importance of dataset partitioning and crossvalidation in the modeling of molecular property prediction models was discussed.An automatic group fragmentation method of ILs is proposed and the effect of group occurrence threshold(evaluated by the number of ILs containing the group in the dataset)on the prediction accuracy is investigated.By comparing five regression algorithms(multiple linear regression,ridge regression,random forest,support vector machine,and neural network),the group contribution model based on ridge regression has the best prediction performance.The average relative error obtained on the composed dataset is 1.88%.
参考文献:
正在载入数据...
