详细信息

Automatic Code Summarization Using Abbreviation Expansion and Subword Segmentation  ( EI收录)  

文献类型:期刊文献

英文题名:Automatic Code Summarization Using Abbreviation Expansion and Subword Segmentation

作者:Liang, Yuguo[1]; Fan, Guisheng[1]; Yu, Huiqun[1]; Li, Mingchen[1]; Huang, Zijie[1]

机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, 200237, China

年份:2024

外文期刊名:SSRN

收录:EI(收录号:20240040598)

语种:英文

外文关键词:Codes (symbols) - Computational linguistics - Deep learning - Inference engines - Learning algorithms - Learning systems - Modeling languages - Natural language processing systems

摘要:Automatic code summarization, the process of automatically generating concise natural language descriptions for code snippets, is critical for enhancing the efficiency of program understanding for software developers and maintainers. Despite the impressive strides made by deep learning-based methods, which have leveraged insights from neural machine translation (NMT) research in the field of natural language processing (NLP), there still exist limitations in their ability of understanding and modeling semantic information due to the unique nature of programming languages. In response, we propose two methods to boost the performance of code summarization models: context-based code abbreviation expansion and unigram language model-based subword segmentation. We employ a series of heuristics to expand abbreviations within identifiers, thereby eliminating the semantic ambiguity associated with these abbreviations and enhancing the language alignment capabilities of code summarization models. Furthermore, we leverage the subword segmentation algorithm to tokenize code into more granular subword sequences, which infuses more semantic information into the training and inference stages of the models, thereby augmenting their program understanding ability. These proposed methods are model-agnostic and can be readily integrated into existing automatic code summarization approaches. Experiments conducted on two widely used Java code summarization datasets demonstrated the effectiveness of these methods. Specifically, by fusing representations of both original and modified codes into the prevailing Transformer model, our presented Semantic Enhanced Transformer for Code Summarization (SETCS) is capable of serving as a robust baseline at the semantic level. Notably, by simply modifying the datasets, our methods achieved performance improvements of up to 7.3%, 10.0%, and 6.7% for representative code summarization models in terms of BLEU-4, METEOR, and ROUGE-L, respectively. ? 2024, The Authors. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心