Lymph Diseases Prediction Using Random Forest and Particle Swarm Optimization — Oak Academic Publishing
Research ArticleOpen AccessGoogle Scholar indexed
Lymph Diseases Prediction Using Random Forest and Particle Swarm Optimization
Department of Computer Science and Information Systems, College of Business Studies, Public Authority for Applied Education and Training, Kuwait City, Kuwait
1 Department of Computer Science and Information Systems, College of Business Studies, Public Authority for Applied Education and Training, Kuwait City, Kuwait
This research aims to develop a model to enhance lymphatic diseases diagnosis by the use of random forest ensemble machine-learning method trained with a simple sampling scheme. This study has been carried out in two major phases: feature selection and classification. In the first stage, a number of discriminative features out of 18 were selected using PSO and several feature selection techniques to reduce the features dimension. In the second stage, we applied the random forest ensemble classification scheme to diagnose lymphatic diseases. While making experiments with the selected features, we used original and resampled distributions of the dataset to train random forest classifier. Experimental results demonstrate that the proposed method achieves a remark-able improvement in classification accuracy rate.
KeywordsClassificationRandom Forest EnsemblePSOSimple Random SamplingInformation Gain RatioSymmetrical Uncertainty
Ciosa, K.J. and Mooree, G.W. (2002) Uniqueness o Medical Data Mining. Artificial Intelligence in Medicine, 26, l-24. http://dx.doi.org/10.1016/s0933-3657(02)00049-0
Ceusters, W. (2000) Medical Natural Language Understanding as a Supporting Technology for Data Mining in Healthcare Medical Data Mining and Knowledge Discovery. Cios KJ Editor, Heidelberg: Springer, pp. 32-60,.
Calle-Alonso, F., Pérez, C.J., Arias-Nicolás, J.P. and Martín, J. (2012) Computer-Aided Diagnosis System: A Bayesian Hybrid Classification Method. Computer Methods and Programs in Biomedicine, 112, 104-113. http://dx.doi.org/10.1016/j.cmpb.2013.05.029
Cselényi, Z. (2005) Mapping the Dimensionality Density and Topology of Data: The Growing Adaptive Neural Gas. Computer Methods and Programs in Biomedicine, 78, 141-156. http://dx.doi.org/10.1016/j.cmpb.2005.02.001
Huang, S.H., Wulsin, L.R., Li, H. and Guo, J. (2009) Dimensionality Reduction for Knowledge Discovery in Medical Claims Database: Application to Antidepressant Medication Utilization Study. Computer Methods and Programs in Biomedicine, 93, 115-123. http://dx.doi.org/10.1016/j.cmpb.2008.08.002
Luukka, P. (2011) Feature Selection Using Fuzzy Entropy Measures with Similarity Classifier. Expert Systems with Applications, 38, 4600-4607. http://dx.doi.org/10.1016/j.eswa.2010.09.133
Luciani, A., Itti, E., Rahmouni, A., Michel Meignan, M. and Clement, O. (2006) Lymph Node Imaging: Basic Principles. European Journal of Radiology, 58, 338-344. http://dx.doi.org/10.1016/j.ejrad.2005.12.038
Sharma, R., Wendt, J.A., Rasmussen, J.C., Adams, A.E., Marshall, M.V. and Sevick-Muraca, E.M. (2008) New Horizons for Imaging Lymphatic Function. Annals of the New York Academy of Sciences, 1131, 13-36. http://dx.doi.org/10.1196/annals.1413.002
Guermazi, A., Brice, P., Hennequin, C. and Sarfati, E. (2003) Lymphography: An Old Technique Retains Its Usefulness. RadioGraphics, 23, 1541-1558. http://dx.doi.org/10.1148/rg.236035704
Cancer Research UK. http://www.cancerresearchuk.org
Karabulut, E.M., Özel, S.A. and Íbrikci, T. (2012) A Comparative Study on the Effect of Feature Selection on Classification Accuracy. Procedia Technology, 1, 323-327. http://dx.doi.org/10.1016/j.protcy.2012.02.068
Derrac, J., Cornelis, C., García, S. and Herrera, F. (2012) Enhancing Evolutionary Instance Selection Algorithms by Means of Fuzzy Rough Set Based Feature Selection. Information Sciences, 186, 73-92. http://dx.doi.org/10.1016/j.ins.2011.09.027
Madden, M.G. (2009) On the Classification Performance of TAN and General Bayesian Networks. Knowledge-Based Systems, 22, 489-495. http://dx.doi.org/10.1016/j.knosys.2008.10.006
De Falco, I. (2013) Differential Evolution for Automatic Rule Extraction from Medical Databases. Applied Soft Computing, 13, 1265-1283. http://dx.doi.org/10.1016/j.asoc.2012.10.022
Abellán, J. and Masegosa, A.R. (2012) Bagging Schemes on the Presence of Class Noise in Classification. Expert Systems with Applications, 39, 6827-6837. http://dx.doi.org/10.1016/j.eswa.2012.01.013
Kennedy, J. and Eberhart, R.C. (2001) Swarm Intelligence. Morgan Kaufmann Publishers, Burlington.
Kennedy, J. and Eberhart, R.C. (1997) A Discrete Binary Version of the Particle Swarm Algorithm. Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics, 5, 4104-4108. http://dx.doi.org/10.1109/icsmc.1997.637339
Quinlan, J.R. (1993) C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, Burlington.
Fayyad, U. and Irani, K. (1993) Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning. Proceedings of the 13th International Joint Conference on Artificial Intelligence, Chambéry, 28 August-3 September 1993, 1022-1027.
Liu, H., Hussain, F., Tan, C. and Dash, M. (2002) Discretization: An Enabling Technique. Data Mining and Knowledge Discovery, 6, 393-423. http://dx.doi.org/10.1023/A:1016304305535
Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32.
Dietterich, T.G. (2000) An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization. Machine Learning, 40, 139-157. http://dx.doi.org/10.1023/A:1007607513941
Breiman, L. (1996) Bagging Predictors. Machine Learning, 24, 123-140.
Weiss, G. and Provost, F. (2003) Learning when Training Data Are Costly: The Effect of Class Distribution on Tree Induction. Journal of Artificial Intelligence Research, 19, 315-354.
Park, B.-H., Ostrouchov, G., Samatova, N.F. and Geist, A. (2004) Reservoir-Based Random Sampling with Replacement from Data Stream. Proceedings of the 2004 SIAM International Conference on Data Mining, 22-24 April 2004, Lake Buena Vista, 492-496. http://dx.doi.org/10.1137/1.9781611972740.53
Mitra, S.K. and Pathak, P.K. (1984) The Nature of Simple Random Sampling. The Annals of Statistics, 12, 1536-1542. http://dx.doi.org/10.1214/aos/1176346810
Schumacher, M., Hollander, N. and Sauerbrei, W. (1997) Resampling and Cross-Validation Techniques: A Tool to Reduce bias Caused by Model Building? Statistics in Medicine, 16, 2813-2827. http://dx.doi.org/10.1002/(SICI)1097-0258(19971230)16:24 3.0.CO;2-Z
Ben-David, A. (2008) Comparison of Classification Accuracy Using Cohen’s Weighted Kappa. Expert Systems with Applications, 34, 825-832. http://dx.doi.org/10.1016/j.eswa.2006.10.022
Cestnik, G., Konenenko, I. and Bratko, I. (1987) Assistant-86: A Knowledge-Elicitation Tool for Sophisticated Users. In: Bratko, I. and Lavrac, N., Eds., Progress in Machine Learning, Sigma Press, Wilmslow, 31-45.