详细信息

IMCStacking: Cost-sensitive stacking learning with feature inverse mapping for imbalanced problems  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:IMCStacking: Cost-sensitive stacking learning with feature inverse mapping for imbalanced problems

作者:Cao, Chenjie[1,2];Wang, Zhe[1,2]

机构:[1]East China Univ Sci & Technol, Minist Educ, Key Lab Adv Control & Optimizat Chem Proc, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2018

卷号:150

起止页码:27

外文期刊名:KNOWLEDGE-BASED SYSTEMS

收录:;EI(收录号:20181705054576);WOS:【SCI-EXPANDED(收录号:WOS:000433654400003)】;

基金:This work is supported by Natural Science Foundations of China under Grant No.61672227, "Shuguang Program" supported by Shanghai Education Development Foundation and Shanghai Municipal Education Commission, and "Action Plan for Innovation on Science and Technology" Projects of Shanghai under Grant No.16511101000.

语种:英文

外文关键词:Feature mapping; Cost-sensitive; Decision tree ensemble; Linear classifiers; Stacking; Imbalanced problems

摘要:Stacking related methods develop rapidly recent years. However, few Stacking based ensemble methods are designed for imbalanced problems. In this paper, a novel Feature Inverse Mapping based Cost sensitive Stacking learning (IMCStacking) is proposed to solve the problems encountered in imbalanced classification. In IMCStacking, we integrate the cost-sensitive Logistic Regression as the final classifier to regard different costs to majority and minority samples. Furthermore, a quick and effective feature inverse mapping technique is applied to IMCStacking to maximize the utilization of the cross-validation process during the Stacking ensemble. This trick can make the proposed method learn better classification thresholds for imbalanced problems. As the result, IMCStacking implements the cost-sensitive strategy on both data level and feature level to overcome the imbalances. Moreover, both linear and forest based approaches work as base classifiers in IMCStacking to guarantee enough generalization. Finally, comprehensive comparison experiments about training times and mean accuracy (M-ACC) on typical imbalanced datasets from KEEL demonstrate both the effectiveness and efficiency of the proposed IMCStacking. (C) 2018 Elsevier B.V. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心