Ensemble-based active learning for class imbalance problem
- 1
- 2
Abstract
In medical diagnosis, the problem of class imbalance is popular. Though there are abundant unlabeled data, it is very difficult and expensive to get labeled ones. In this paper, an ensemble-based active learning algorithm is proposed to address the class imbalance problem. The artificial data are created according to the distribution of the training dataset to make the ensemble diverse, and the random subspace re-sampling method is used to reduce the data dimension. In selecting member classifiers based on misclassification cost estimation, the minority class is assigned with higher weights for misclassification costs, while each testing sample has a variable penalty factor to induce the ensemble to correct current error. In our experiments with UCI disease datasets, instead of classification accuracy, F-value and G-means are used as the evaluation rule. Compared with other ensemble methods, our method shows best performance, and needs less labeled samples.
- Japkowicz, N. and Stephen, S. (2002) The class imbalance problem: A systematic study. Intelligent Data Analysis, 6(5), 203-231.
- Gustavo, E.A., Batista, P.A., Ronaldo, C., et a1. (2004) A study of the behavior of several methods for balancing machine learning training data. SIGKDD Explorations, 6(1), 20-29.
- Settles, B. (2009) Active Learning Literature Survey. Computer Sciences Technical Report 1648, University of Wisconsin-Madison.
- Tomek, I. (1976) Two modifications of CNN. IEEE Transaction on Systems Man and Communications, 6, 769-772.
- Hart, P.E. (1968) The condensed nearest neighbor rule. IEEE Transaction on Information Theory, 14(3), 515- 516.
- Laurikkala, J. (2001) Improving identification of difficult small classes by balancing class distribution. Proceedings of the 8th Conference on AI in Medicine, Cascais, Portugal, Europe: Artificial Intelligence Medicine, 63-66.
- Wilson, D.L. (1972) Asymptotic properties of nearest neighbor rules using edited data.IEEE Transaction on Systems, Man and Communications, 2(3), 408-421.
- Chawlan, V., Bowyer, K.W. and Hall, L.O. (2002) SMOTE: Synthetie minority over-sampling technique. Journal of Aflificial Intelligence Research, 16(1), 321- 357.
- Joshi, M., Kumar, V. and Agarwal, R. (2001) Evaluating boosting algorithms to classify rare classes: Comparison and improvements. Proceedings of the 1st IEEE Interna- tional Conference on Data Mining. Washington DC: IEEE Computer Society, 257-264.
- Akbani, R., Kwek, S. and Japkowicz, N. (2004) Applying support vector machines to imbalanced datasets. Proceedings of the 15th European Conference on Machines Learning, Pisa, Italy, 39-50.
- Krogh, A. and Vedelsby, J. (1995) Neural network ensembles, cross validation and active learning. Advances in Neural Information Processing Systems, 7, 231-238.
- Provost, F. (2000) Machine learning from imbalanced data sets 101. Invited paper for the AAAI, Workshop on Imbalanced Data Sets, Menlo Park, CA.
- Abe, N. (2003) Invited talk: Sampling approaches to learning from imbalanced datasets: Active learning, cost sensitive learning and beyond. ICML-KDD Workshop: Learning from Imbalanced Data Sets.
- Ertekin, S., Huang, J. and Giles, C.L. (2007) Active learning for class imbalance problem. Proceedings of Annual International ACM SIGIR Conference Research and development in information retrieval, Amsterdam, Netherlands, 823-824.