详细信息
Partition dataset according to amino acid type improves the prediction of deleterious non-synonymous SNPs ( SCI-EXPANDED收录)
文献类型:期刊文献
英文题名:Partition dataset according to amino acid type improves the prediction of deleterious non-synonymous SNPs
作者:Yang, Jing[1,2];Li, Yuan-Yuan[1,2];Li, Yi-Xue[1,2];Ye, Zhi-Qiang[3,4]
机构:[1]E China Univ Sci & Technol, Sch Biotechnol, Shanghai 200237, Peoples R China;[2]Shanghai Ctr Bioinformat Technol, Shanghai 200235, Peoples R China;[3]Peking Univ, Shenzhen Grad Sch, Sch Chem Biol & Biotechnol, Lab Chem Genom, Shenzhen 518055, Peoples R China;[4]Chinese Acad Sci, Shanghai Inst Biol Sci, Key Lab Syst Biol, Shanghai 200031, Peoples R China
年份:2012
卷号:419
期号:1
起止页码:99
外文期刊名:BIOCHEMICAL AND BIOPHYSICAL RESEARCH COMMUNICATIONS
收录:;WOS:【SCI-EXPANDED(收录号:WOS:000301560500018)】;
基金:This work was supported by grants from the National "973" Basic Research Program of China (2012CB316501), the National '863' Hi-Tech Research and Development Program of China (2009AA022710), the National Natural Science Foundation of China (30800641, 31171268, 31000380, 30900834). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. We thank Hui Yu, Rudong Li, Fudong Yu, Yunqin Chen, and Baohong Liu for helpful discussions.
语种:英文
外文关键词:Non-synonymous SNP; Disease-association; Dataset partition; Machine learning
摘要:Many non-synonymous SNPs (nsSNPs) are associated with diseases, and numerous machine learning methods have been applied to train classifiers for sorting disease-associated nsSNPs from neutral ones. The continuously accumulated nsSNP data allows us to further explore better prediction approaches. In this work, we partitioned the training data into 20 subsets according to either original or substituted amino acid type at the nsSNP site. Using support vector machine (SVM), training classification models on each subset resulted in an overall accuracy of 76.3% or 74.9% depending on the two different partition criteria, while training on the whole dataset obtained an accuracy of only 72.6%. Moreover, the dataset was also randomly divided into 20 subsets, but the corresponding accuracy was only 73.2%. Our results demonstrated that partitioning the whole training dataset into subsets properly, i.e., according to the residue type at the nsSNP site, will improve the performance of the trained classifiers significantly, which should be valuable in developing better tools for predicting the disease-association of nsSNPs. (C) 2012 Elsevier Inc. All rights reserved.
参考文献:
正在载入数据...
