详细信息
Mining K-mers of Various Lengths in Biological Sequences ( CPCI-S收录 EI收录)
文献类型:会议论文
英文题名:Mining K-mers of Various Lengths in Biological Sequences
作者:Zhang, Jingsong[1];Guo, Jianmei[2];Yu, Xiaoqing[3];Yu, Xiangtian[1];Guo, Weifeng[4];Zeng, Tao[1];Chen, Luonan[1,5]
机构:[1]Chinese Acad Sci, Shanghai Inst Biol Sci, Inst Biochem & Cell Biol, Shanghai 200031, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[3]Shanghai Inst Technol, Dept Appl Math, Shanghai 201418, Peoples R China;[4]Northwestern Polytech Univ, Sch Automat, Xian 710072, Shaanxi, Peoples R China;[5]Univ Tokyo, Collaborat Res Ctr Innovat Math Modelling, Inst Ind Sci, Tokyo 1538505, Japan
会议论文集:13th International Symposium on Bioinformatics Research and Applications (ISBRA)
会议日期:MAY 29-JUN 02, 2017
会议地点:Honolulu, HI
语种:英文
外文关键词:Sequential pattern mining; K-mer counting; K-mers of various lengths; Biological sequence analysis
摘要:Counting the occurrence frequency of each k-mer in a biological sequence is an important step in many bioinformatics applications. However, most k-mer counting algorithms rely on a given k to produce single-length k-mers, which is inefficient for sequence analysis for different k. Moreover, existing k-mer counters focus more on DNA sequences and less on protein ones. In practice, the analysis of k-mers in protein sequences can provide substantial biological insights in structure, function and evolution. To this end, an efficient algorithm, called VLmer (Various Length k-mer mining), is proposed to mine k-mers of various lengths termed vl-mers via inverted-index technique, which is orders of magnitude faster than the conventional forward-index method. Moreover, to the best of our knowledge, VLmer is the first able to mine k-mers of various lengths in both DNA and protein sequences.
参考文献:
正在载入数据...
