详细信息

Sample and feature selecting based ensemble learning for imbalanced problems  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Sample and feature selecting based ensemble learning for imbalanced problems

作者:Wang, Zhe[1,2];Jia, Peng[2];Xu, Xinlei[2];Wang, Bolu[2];Zhu, Yujin[2];Li, Dongdong[1,2]

机构:[1]East China Univ Sci & Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2021

卷号:113

外文期刊名:APPLIED SOFT COMPUTING

收录:;EI(收录号:20213810906066);WOS:【SCI-EXPANDED(收录号:WOS:000710723900012)】;

基金:This work is supported by Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research"under Grant No. 20511100600, Shanghai Science and Technology Program "Federated based cross-domain and cross-task incremental learning"under Grant No. 21511100800, Natural Science Foundation of China under Grant No. 62076094, National Key Research and Development Project of Ministry of Science and Technology of China under Grant No. 2018AAA 0101302, Natural Science Foundations of China under Grant No. 61806078.

语种:英文

外文关键词:Imbalanced classification; Ensemble learning; Random forest; Heart failure mortality prediction

摘要:Imbalanced problem is concerned with the performance of classifiers on the data set with severe class imbalance distribution. Traditional methods are misled by the majority samples to make the incorrect prediction and fail to make full use of minority samples. This paper is motivated to design a novel hybrid ensemble learning strategy named Sample and Feature Selection Hybrid Ensemble Learning (SFSHEL) and combine it with random forest to improve the classification performance of imbalanced data. Specifically, SFSHEL considers cluster-based stratification to undersample the majority samples and adopts sliding windows mechanism to generate a diversity of feature subsets, simultaneously. Then the weights trained through validation are assigned to different base learners and SFSHEL makes the prediction by weighted voting at last. In this manner, SFSHEL could not only guarantee the acceptable performance, but also save computational time. Furthermore, the weighting process makes SFSHEL interpret the importance of each selected feature set, which is important in the real-world scenarios. The contributions of the proposed strategy are: (1) reducing the impact of class imbalance distribution, (2) assigning based learner weights only once after the training process, and (3) generating weights of features to help interpret the importance of clinical features. In practice, the random forest is adopted as the base learner for SFSHEL, so as to build a classifier abbreviated as SFSHEL-RF. The experiments show the average performance of the proposed SFSHEL-RF on a part of KEEL dataset reaches 91.37%, which is comparable to our previous best ECUBoost-RF method and higher than the other eleven methods. On the clinical heart failure datasets, the performance of SFSHEL-RF can stably reach the level of the top three with three indicators. The experimental results on both the standard imbalanced and clinical heart failure datasets validate the effectiveness and stability of SFSHEL-RF. (C) 2021 Elsevier B.V. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心