详细信息

Local versus Global Models for Just-In-Time Software Defect Prediction  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Local versus Global Models for Just-In-Time Software Defect Prediction

作者:Yang, Xingguang[1,2];Yu, Huiqun[1,3];Fan, Guisheng[1];Shi, Kai[1];Chen, Liqiong[4]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]Shanghai Key Lab Comp Software Evaluating & Testi, Shanghai 201112, Peoples R China;[3]Shanghai Engn Res Ctr Smart Energy, Shanghai, Peoples R China;[4]Shanghai Inst Technol, Dept Comp Sci & Informat Engn, Shanghai 201418, Peoples R China

年份:2019

卷号:2019

外文期刊名:SCIENTIFIC PROGRAMMING

收录:;EI(收录号:20192707139541);WOS:【SCI-EXPANDED(收录号:WOS:000472887800001)】;

基金:This work was partially supported by the NSF of China under Grant nos. 61772200 and 61702334, Shanghai Pujiang Talent Program under Grant no. 17PJ1401900, Shanghai Municipal Natural Science Foundation under Grant nos. 17ZR1406900 and 17ZR1429700, Educational Research Fund of ECUST under Grant no. ZH1726108, and Collaborative Innovation Foundation of Shanghai Institute of Technology under Grant no. XTCX2016-20.

语种:英文

外文关键词:Forecasting - Just in time production - Open source software

摘要:Just-in-time software defect prediction (JIT-SDP) is an active topic in software defect prediction, which aims to identify defect-inducing changes. Recently, some studies have found that the variability of defect data sets can affect the performance of defect predictors. By using local models, it can help improve the performance of prediction models. However, previous studies have focused on module-level defect prediction. Whether local models are still valid in the context of JIT-SDP is an important issue. To this end, we compare the performance of local and global models through a large-scale empirical study based on six open-source projects with 227417 changes. The experiment considers three evaluation scenarios of cross-validation, cross-project-validation, and timewise-cross-validation. To build local models, the experiment uses the k-medoids to divide the training set into several homogeneous regions. In addition, logistic regression and effort-aware linear regression (EALR) are used to build classification models and effort-aware prediction models, respectively. The empirical results show that local models perform worse than global models in the classification performance. However, local models have significantly better effort-aware prediction performance than global models in the cross-validation and cross-project-validation scenarios. Particularly, when the number of clusters k is set to 2, local models can obtain optimal effort-aware prediction performance. Therefore, local models are promising for effort-aware JIT-SDP.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心