TY - GEN
T1 - An imbalanced data classification method based on automatic clustering under-sampling
AU - Deng, Xiaoheng
AU - Zhong, Weijian
AU - Ren, Ju
AU - Zeng, Detian
AU - Zhang, Honggang
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2017/1/17
Y1 - 2017/1/17
N2 - Classification of imbalanced datasets has become one of the most challenging problems in big data mining. Because the number of positive samples is far less than the negative samples, low accuracy and poor generalization performance and some other defects always go with learning process of traditional algorithms. Ensemble construction algorithm is an important method to handle this problem. Especially, the ensemble construction algorithm based on random under-sampling or clustering can effectively improve the performance of classification. However, the former causes information loss easily and the latter increases complexity. In this paper, we propose ACUS, an improved ensemble algorithm based on automatic clustering and under-sampling. ACUS conducts clustering first according to the weight of samples, and then it constructs balanced-distributed dataset which consists of a certain percentage of the majority class and all of the minority class from each cluster. With Adaboost algorithm construction, these datasets are used to get an ensemble classifier. Experimental results demonstrate the advantages of our proposed algorithm in terms of accuracy, simplicity and high stability.
AB - Classification of imbalanced datasets has become one of the most challenging problems in big data mining. Because the number of positive samples is far less than the negative samples, low accuracy and poor generalization performance and some other defects always go with learning process of traditional algorithms. Ensemble construction algorithm is an important method to handle this problem. Especially, the ensemble construction algorithm based on random under-sampling or clustering can effectively improve the performance of classification. However, the former causes information loss easily and the latter increases complexity. In this paper, we propose ACUS, an improved ensemble algorithm based on automatic clustering and under-sampling. ACUS conducts clustering first according to the weight of samples, and then it constructs balanced-distributed dataset which consists of a certain percentage of the majority class and all of the minority class from each cluster. With Adaboost algorithm construction, these datasets are used to get an ensemble classifier. Experimental results demonstrate the advantages of our proposed algorithm in terms of accuracy, simplicity and high stability.
KW - Boosting
KW - Class distribution
KW - Classification
KW - Ensemble
KW - Imbalanced datasets
UR - https://www.scopus.com/pages/publications/85013409014
UR - https://www.scopus.com/pages/publications/85013409014#tab=citedBy
U2 - 10.1109/PCCC.2016.7820640
DO - 10.1109/PCCC.2016.7820640
M3 - Conference contribution
AN - SCOPUS:85013409014
T3 - 2016 IEEE 35th International Performance Computing and Communications Conference, IPCCC 2016
BT - 2016 IEEE 35th International Performance Computing and Communications Conference, IPCCC 2016
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 35th IEEE International Performance Computing and Communications Conference, IPCCC 2016
Y2 - 9 December 2016 through 11 December 2016
ER -