详细信息
Cancer classification based on chromatin accessibility profiles with deep adversarial learning model ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Cancer classification based on chromatin accessibility profiles with deep adversarial learning model
作者:Yang, Hai[1];Wei, Qiang[2,3];Li, Dongdong[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai, Peoples R China;[2]Vanderbilt Univ, Dept Mol Physiol & Biophys, Nashville, TN 37232 USA;[3]Vanderbilt Univ, Vanderbilt Genet Inst, 221 Kirkland Hall, Nashville, TN 37235 USA
年份:2020
卷号:16
期号:11
外文期刊名:PLOS COMPUTATIONAL BIOLOGY
收录:;EI(收录号:20223012429092);WOS:【SCI-EXPANDED(收录号:WOS:000591834800004)】;
基金:DL was supported by the National Major Scientific and Technological Special Project for "Significant New Drugs Development" under Grant No. 2019ZX09201004, ZW was supported by "Shuguang Program" supported by the Shanghai Education Development Foundation and Shanghai Municipal Education Commission. HY was supported by the Natural Science Foundation of China under Grant No. 61902126. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
语种:英文
外文关键词:Classification (of information) - Deep learning - Diagnosis - Diseases - Genes
摘要:Author summary A few methods have been developed to leverage omics data (e.g., iCluster) to solve cancer classification. However, these omics data always focus on the coding regions (2% of the human genome). Cancer classification methods based on the high-dimensional raw data across the whole genome are rare. Our approach addressed crucial and fundamental limitations in existing approaches. ClusterATAC used adversarial learning to handle the limited but high-dimensional omics data. Feature selection techniques are used to analyze the essential non-coding loci and their regulatory genes for each cluster. The outcome can lead to a deeper understanding of the regulatory mechanisms that lead to cancer development and progression. We successfully obtained 22 cancer subgroups form the ATAC-seq profiles of 401 TCGA samples. We observed that most subgroups follow the 'Cell-of-Origin' pattern, consistent with the recent study. There were significant differences between the 22 clusters. More than 70% of the subgroups were homogeneous for a single cancer type. We identified the representative loci and the corresponding regulatory genes of each cluster. We found that these loci and genes are always tumor-specific and responsible for the occurrence and development of the related tumor. Given the complexity and diversity of the cancer genomics profiles, it is challenging to identify distinct clusters from different cancer types. Numerous analyses have been conducted for this propose. Still, the methods they used always do not directly support the high-dimensional omics data across the whole genome (Such as ATAC-seq profiles). In this study, based on the deep adversarial learning, we present an end-to-end approach ClusterATAC to leverage high-dimensional features and explore the classification results. On the ATAC-seq dataset and RNA-seq dataset, ClusterATAC has achieved excellent performance. Since ATAC-seq data plays a crucial role in the study of the effects of non-coding regions on the molecular classification of cancers, we explore the clustering solution obtained by ClusterATAC on the pan-cancer ATAC dataset. In this solution, more than 70% of the clustering are single-tumor-type-dominant, and the vast majority of the remaining clusters are associated with similar tumor types. We explore the representative non-coding loci and their linked genes of each cluster and verify some results by the literature search. These results suggest that a large number of non-coding loci affect the development and progression of cancer through its linked genes, which can potentially advance cancer diagnosis and therapy.
参考文献:
正在载入数据...
