详细信息
非相关线性判别分析用于蛋白质组数据的分类及特征挑选(英文)
Classification and feature selection of proteomic data by uncorrelated linear discriminant analysis
文献类型:期刊文献
中文题名:非相关线性判别分析用于蛋白质组数据的分类及特征挑选(英文)
英文题名:Classification and feature selection of proteomic data by uncorrelated linear discriminant analysis
作者:张明锦[1,2];王文明[1];童佩瑾[1];杜一平[1]
机构:[1]华东理工大学分析测试中心,上海200237;[2]青海师范大学化学系,青海西宁810008
年份:2009
卷号:26
期号:12
起止页码:1563
中文期刊名:计算机与应用化学
外文期刊名:Computers and Applied Chemistry
收录:CSTPCD;;北大核心:【北大核心2008】;CSCD:【CSCD2011_2012】;
基金:国家自然科学基金(20975039)资助项目
语种:中文
中文关键词:非相关线性判别分析;卡方检验;蛋白质组学;质谱数据;生物标记物
外文关键词:uncorrelated linear discriminant analysis, Chi-squared test, proteomics, MS data, biomarker
摘要:提出一种非相关线性判别分析(ULDA)结合统计卡方检验(CHI2)的方法用于蛋白质组质谱数据的分类及特征挑选.首先以卡方检验为过滤器去除无类间差别的变量,然后用ULDA进行样本分类与特征筛选,通过对两组数据的分析,最终选择出的特征变量在这两组数据中的特异性分别为98.2%和95.74%,灵敏度均为100%.结果表明本文提出的方法能较好地处理变量数很大的蛋白质组数据,同时表明最后选择的特征变量有可能作为潜在的生物标记物,为相关疾病的早期诊断提供线索.
A uncorrelated linear discriminant analysis (ULDA) combined with Chi-squared (CHI2) method was proposed in this paper and was used to classification and feature selection for proteomic MS data. The method uses CHI2 method as a filter for eliminates the irrelative variables for classification firstly, and then performs ULDA for sample classification and feature selection. After analysis for 2 datasets, the selected variables obtained 98.2% and 95.74% specificity respectively, and 100% sensitivity for both. It can be inferred from the results that it is possible to differentiate between control and cancer samples using the proposed approach, it is also possible that the selected variables can be regard as potential biomarkers that provide clues for disease earlier detection.
参考文献:
正在载入数据...
