详细信息

An improved sequence based prediction protocol for DNA-binding proteins using SVM and comprehensive feature analysis  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:An improved sequence based prediction protocol for DNA-binding proteins using SVM and comprehensive feature analysis

作者:Zou, Chuanxin[1];Gong, Jiayu[1];Li, Honglin[1]

机构:[1]E China Univ Sci & Technol, Shanghai Key Lab New Drug Design, State Key Lab Bioreactor Engn, Sch Pharm, Shanghai 200237, Peoples R China

年份:2013

卷号:14

外文期刊名:BMC BIOINFORMATICS

收录:;EI(收录号:20131316153006);WOS:【SCI-EXPANDED(收录号:WOS:000316396500001)】;

基金:This work was supported by the Fundamental Research Funds for the Central Universities, the National Natural Science Foundation of China (grants 21173076, 81102375, 81230090, 81222046 and 81230076), the Special Fund for Major State Basic Research Project (grant 2009CB918501), the Shanghai Committee of Science and Technology (grant 11DZ2260600), and the 863 Hi-Tech Program of China (grant 2012AA020308). Honglin Li is also sponsored by Program for New Century Excellent Talents in University (grant NCET-10-0378).

语种:英文

外文关键词:Forecasting - Feature Selection - DNA

摘要:Background: DNA-binding proteins (DNA-BPs) play a pivotal role in both eukaryotic and prokaryotic proteomes. There have been several computational methods proposed in the literature to deal with the DNA-BPs, many informative features and properties were used and proved to have significant impact on this problem. However the ultimate goal of Bioinformatics is to be able to predict the DNA-BPs directly from primary sequence. Results: In this work, the focus is how to transform these informative features into uniform numeric representation appropriately and improve the prediction accuracy of our SVM-based classifier for DNA-BPs. A systematic representation of some selected features known to perform well is investigated here. Firstly, four kinds of protein properties are obtained and used to describe the protein sequence. Secondly, three different feature transformation methods (OCTD, AC and SAA) are adopted to obtain numeric feature vectors from three main levels: Global, Nonlocal and Local of protein sequence and their performances are exhaustively investigated. At last, the mRMR-IFS feature selection method and ensemble learning approach are utilized to determine the best prediction model. Besides, the optimal features selected by mRMR-IFS are illustrated based on the observed results which may provide useful insights for revealing the mechanisms of protein-DNA interactions. For five-fold cross-validation over the DNAdset and DNAaset, we obtained an overall accuracy of 0.940 and 0.811, MCC of 0.881 and 0.614 respectively. Conclusions: The good results suggest that it can efficiently develop an entirely sequence-based protocol that transforms and integrates informative features from different scales used by SVM to predict DNA-BPs accurately. Moreover, a novel systematic framework for sequence descriptor-based protein function prediction is proposed here.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心