详细信息

Mining K-mers of Various Lengths in Biological Sequences  ( CPCI-S收录 EI收录)  

文献类型:会议论文

英文题名:Mining K-mers of Various Lengths in Biological Sequences

作者:Zhang, Jingsong[1];Guo, Jianmei[2];Yu, Xiaoqing[3];Yu, Xiangtian[1];Guo, Weifeng[4];Zeng, Tao[1];Chen, Luonan[1,5]

机构:[1]Chinese Acad Sci, Shanghai Inst Biol Sci, Inst Biochem & Cell Biol, Shanghai 200031, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[3]Shanghai Inst Technol, Dept Appl Math, Shanghai 201418, Peoples R China;[4]Northwestern Polytech Univ, Sch Automat, Xian 710072, Shaanxi, Peoples R China;[5]Univ Tokyo, Collaborat Res Ctr Innovat Math Modelling, Inst Ind Sci, Tokyo 1538505, Japan

会议论文集:13th International Symposium on Bioinformatics Research and Applications (ISBRA)

会议日期:MAY 29-JUN 02, 2017

会议地点:Honolulu, HI

语种:英文

外文关键词:Sequential pattern mining; K-mer counting; K-mers of various lengths; Biological sequence analysis

摘要:Counting the occurrence frequency of each k-mer in a biological sequence is an important step in many bioinformatics applications. However, most k-mer counting algorithms rely on a given k to produce single-length k-mers, which is inefficient for sequence analysis for different k. Moreover, existing k-mer counters focus more on DNA sequences and less on protein ones. In practice, the analysis of k-mers in protein sequences can provide substantial biological insights in structure, function and evolution. To this end, an efficient algorithm, called VLmer (Various Length k-mer mining), is proposed to mine k-mers of various lengths termed vl-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index method. Moreover, to the best of our knowledge, VLmer is the first able to mine k-mers of various lengths in both DNA and protein sequences.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心