MultiDMet: Designing a Hybrid Multidimensional Metrics Framework to Predictive Modeling for Performance Evaluation and Feature Selection — Oak Academic Publishing
Research ArticleOpen AccessGoogle Scholar indexed
MultiDMet: Designing a Hybrid Multidimensional Metrics Framework to Predictive Modeling for Performance Evaluation and Feature Selection
Department of Information and Communication Engineering, Addis Ababa Science and Technology University, Addis Ababa, Ethiopia
,
Department of Computer Science, HILCOE School of Computer Science and Technology, Addis Ababa, Ethiopia
1 Department of Information and Communication Engineering, Addis Ababa Science and Technology University, Addis Ababa, Ethiopia
2 Department of Computer Science, HILCOE School of Computer Science and Technology, Addis Ababa, Ethiopia
In a competitive digital age where data volumes are increasing with time, the ability to extract meaningful knowledge from high-dimensional data using machine learning (ML) and data mining (DM) techniques and making decisions based on the extracted knowledge is becoming increasingly important in all business domains. Nevertheless, high-dimensional data remains a major challenge for classification algorithms due to its high computational cost and storage requirements. The 2016 Demographic and Health Survey of Ethiopia (EDHS 2016) used as the data source for this study which is publicly available contains several features that may not be relevant to the prediction task. In this paper, we developed a hybrid multidimensional metrics framework for predictive modeling for both model performance evaluation and feature selection to overcome the feature selection challenges and select the best model among the available models in DM and ML. The proposed hybrid metrics were used to measure the efficiency of the predictive models. Experimental results show that the decision tree algorithm is the most efficient model. The higher score of HMM ( m, r ) = 0.47 illustrates the overall significant model that encompasses almost all the user’s requirements, unlike the classical metrics that use a criterion to select the most appropriate model. On the other hand, the ANNs were found to be the most computationally intensive for our prediction task. Moreover, the type of data and the class size of the dataset (unbalanced data) have a significant impact on the efficiency of the model, especially on the computational cost, and the interpretability of the parameters of the model would be hampered. And the efficiency of the predictive model could be improved with other feature selection algorithms (especially hybrid metrics) considering the experts of the knowledge domain, as the understanding of the business domain has a significant impact.
Molina, L.C., Belanche, L. and Nebot, A. (2022) Feature Selection Algorithms: A Survey and Experimental Evaluation. 2002 IEEE International Conference on Data Mining, Maebashi City, 9-12 December 2002, 306-313.
Liu, H., Motoda, H. and Yu, L. (2004) Selective Sampling Approach to Active Feature Selection. Artificial Intelligence, 159, 49-74. https://doi.org/10.1016/j.artint.2004.05.009
Chandrashekar, G. and Sahin, F. (2014) A Survey on Feature Selection Methods. Computers & Electrical Engineering, 40, 16-28. https://doi.org/10.1016/j.compeleceng.2013.11.024
Nakariyakul, S. (2014) Suboptimal Branch and Bound Algorithms for Feature Subset Selection: A Comparative Study. Computers & Electrical Engineering, 45, 62-70. https://doi.org/10.1016/j.patrec.2014.03.002
Sheikhpour, R., Sarram, M.A., Gharaghani, S. and Chahooki, M.A.Z. (2017) A Survey on Semi-Supervised Feature Selection Methods. Pattern Recognition, 64, 141-158. https://doi.org/10.1016/j.patcog.2016.11.003
Agrawal, R. and Psaila, G. (1995) Active Data Mining. Proceedings of the First International Conference on Knowledge Discovery and Data Mining, Montreal, 20-21 August 1995, 3-8.
Liao, S.H., Chu, P.H. and Hsiao, P.Y. (2012) Data Mining Techniques and Applications—A Decade Review from 2000 to 2011. Expert Systems with Applications, 39, 11303-11311. https://doi.org/10.1016/j.eswa.2012.02.063
Janecek, A.G.K., Gansterer, G.F., et al. (2008) On the Relationship between Feature Selection and Classification Accuracy. New challenges for feature selection in data mining and knowledge discovery, Antwerp, 15 September 2008, 90-105.
Bellman, R. (1961) Adaptive Control Processes: A Guided Tour. Princeton University Press, Princeton. https://doi.org/10.1515/9781400874668
Chumerin, N. and Van Hulle, M.M. (2006) Comparison of Two Feature Extraction Methods Based on Maximization of Mutual Information. 2006 16th IEEE Signal Processing Society Workshop on Machine Learning for Signal Processing, Maynooth, 6-8 September 2006, 343-348. https://doi.org/10.1109/MLSP.2006.275572
Motoda, H. and Liu, H. (2002) Feature Selection, Extraction and Construction. Springer, New York.
Ladla, L. and Deepa, T. (2011) Feature Selection Methods and Algorithms. International Journal on Computer Science and Engineering (IJCSE), 3, 1787-1797.
(2021) Ethiopia and Demographic and Health Survey 2016 [FR328]. https://dhsprogram.com/
Joshi, S., Deepa Shenoy, P., Vibhudendra Simha, G.G., Venugopal, K.R. and Patnaik, L.M. (2010) Classification of Neurodegenerative Disorders Based on Major Risk Factors EmployingMachine Learning Techniques. International Journal of Engineering and Technology, 2, 350-355. https://doi.org/10.7763/IJET.2010.V2.146
Kunwar, V., Chandel, K., Sabitha, A.S. and Bansal, A. (2016) Chronic Kidney Disease Analysis Using Data Mining Classification Techniques. 2016 6th International Conference—Cloud System and Big Data Engineering (Confluence), Noida, 14-15 January 2016, 300-305. https://doi.org/10.1109/CONFLUENCE.2016.7508132
Bellazzi, R. and Zupan, B. (2008) Predictive Data Mining in Clinical Medicine: Current Issues and Guidelines. International Journal of Medical Informatics, 77, 81-97. https://doi.org/10.1016/j.ijmedinf.2006.11.006
Ge, Z.Q., Song, Z.H., Ding, S.X. and Huang, B. (2017) Data Mining and Analytics in the Process Industry: The Role of Machine Learning. IEEE Access, 5, 20590-20616.
Roberts, A. (2005) AI32: Guide to Weka. http://lia.deis.unibo.it/Courses/SistInt/Lucidi/lab01-small_weka_guide.pdf
Han, J. and Kamber, M. (2006) Data Mining: Concepts and Techniques. 2nd Edition, Morgan Kaufmann Publishers, San Francisco.
Brown, G., Pocock, A., Zhao, M.J. and LujÆn, M. (2012) Conditional Likelihood Maximization: A Unifying Framework for Information Theoretic Feature Selection. Journal of Machine Learning Research, 13, 27-66.
Souza, J. (2004) Feature Selection with a General Hybrid Algorithm. Ph.D. Thesis, University of Ottawa, Ottawa.
Yu, L. and Liu, H. (2004) Efficient Feature Selection via Analysis of Relevance and Redundancy. Journal of Machine Learning Research, 5, 1205-1224.
Kohavi, R. and John, G.H. (1997) Wrappers for Feature Subset Selection. Artificial Intelligence, 97, 273-324. https://doi.org/10.1016/S0004-3702(97)00043-X
Nakariyakul, S., Liu, Z.P. and Chen, L. (2012) Detecting Thermophilic Proteins through Selecting Amino Acid and Dipeptide Composition Features. Amino Acids, 42, 1947-1953. https://doi.org/10.1007/s00726-011-0923-1
Dash, M. and Liu, H. (1997) Feature Selection for Classification. Intelligent Data Analysis, 1, 131-156. https://doi.org/10.3233/IDA-1997-1302
Das, S. (2001) Filters, Wrappers and a Boosting-Based Hybrid for Feature Selection. Proceedings of the Eighteenth International Conference on Machine Learning, San Francisco, 28 June-1 July 2001, 74-81.
Guyon, I., Weston, J., Barnhill, S. and Vapnik, V. (2002) Gene Selection for Cancer Classification Using Support Vector Machines. Machine Learning, 46, 389-422. https://doi.org/10.1023/A:1012487302797
Neumann, J., Schnörr, C. and Steidl, G. (2005) Combined SVM-Based Feature Selection and Classification. Machine Learning, 61, 129-150. https://doi.org/10.1007/s10994-005-1505-9
Guyon, I. and Elisseeff, A. (2003) An Introduction to Variable and Feature Selection. Journal of Machine Learning Research, 3, 1157-1182.
Shafique, U., Majeed, F., Qaiser, H. and Mustafa, I.U. (2015) Data Mining in Healthcare for Heart Diseases. International Journal of Innovation and Applied Studies, 10, 1312-1322.
Soni, J., Ansari, U., Sharma, D. and Soni, S. (2011) Predictive Data Mining for Medical Diagnosis: An Overview of Heart Disease Prediction. International Journal of Computer Applications, 17, 43-48. https://doi.org/10.5120/2237-2860
Koh, H.C. and Tan, G. (2011) Data Mining Applications in Healthcare. Journal of Healthcare Information Management, 19, 64-72.
Mahindrakar, P. and Hanumanthappa, M. (2013) Data Mining in Healthcare: A Survey of Techniques and Algorithms with Its Limitations and Challenges. International Journal of Engineering Research and Applications, 3, 937-941.
Hailu, T. (2015) Comparing Data Mining Techniques in HIV Testing Prediction. Intelligent Information Management, 7, 153-180. https://doi.org/10.4236/iim.2015.73014
Brosette, S.E., Spragre, A.P., Jones, W.T. and Moser, S.A. (2000) A Data Mining System for Infection Control Surveillance. Methods of Information in Medicine, 39, 303-310. https://doi.org/10.1055/s-0038-1634449
Sandhya, J., Deepa Shenoy, P., Venugopal, K.R. and Patnaik, A. (2010) Classification and Treatment of Different Stages of Alzheimer’s Disease Using Various Machine Learning Methods. International Journal of Bioinformatics Research, 2, 44-52. https://doi.org/10.9735/0975-3087.2.1.44-52
Giudici, P. (2003) Applied Data Mining: Statistical Methods for Business and Industry. John Wiley, New York.
Kharya, S. (2012) Using Data Mining Techniques for Diagnosis and Prognosis of Cancer Disease. International Journal of Computer Science, Engineering and Information Technology (IJCSEIT), 2, 55-66. https://doi.org/10.5121/ijcseit.2012.2206
Sundar, N.A., Latha, P.P. and Chandra, M.R. (2012) Performance Analysis of Classification Data Mining Techniques over Heart Disease Database. International Journal of Engineering Science & Advanced Technology, 2, 470-478.
Obenshain, M.K. (2004) Application of Data Mining Techniques to Healthcare Data. Infection Control and Hospital Epidemiology, 25, 690-695. https://doi.org/10.1086/502460
Maniya, H., Hasan, M. and Patel, K.P. (2011) Comparative Study of Naïve Bayes Classifier and KNN for Tuberculosis. International Conference on Web Services Computing (ICWSC), 2, 22-26.
Kusiak, A., Dixon, B. and Shah, S. (2005) Predicting Survival Time for Kidney Dialysis Patients: A Data Mining Approach. Computers in Biology and Medicine, 35, 311-327. https://doi.org/10.1016/j.compbiomed.2004.02.004
Shetty, D., Rit, K., Shaikh, S. and Patil, N. (2017) Diabetes Disease Prediction Using Data Mining. 2017 International Conference on Innovations in Information, Embedded and Communication Systems (ICIIECS), Coimbatore, 17-18 March 2017, 1-5. https://doi.org/10.1109/ICIIECS.2017.8276012
Rahim, N.F., Taib, S.M. and Abidin, A.I.Z. (2017) Dengue Fatality Prediction Using Data Mining. Journal of Fundamental and Applied Sciences, 9, 671-683. https://doi.org/10.4314/jfas.v9i6s.52
Uhmn, S., Kim, D.H., Cho, S.W., Cheong, J.Y. and Kim, J. (2007) Chronic Hepatitis Classification Using SNP Data and Data Mining Techniques. 2007 Frontiers in the Convergence of Bioscience and Information Technologies, Jeju, 11-13 October 2007, 81-86. https://doi.org/10.1109/FBIT.2007.64
Passmore, L., Goodside, J., Hamel, L., Gonzalez, L., Silberstein, T.A.L.I. and Trimarchi, J.A.M.E.S. (2003) Assessing Decision Tree Models for Clinical in-vitro Fertilization Data. Dept. of Computer Science and Statistics University of Rhode Island, Technical Report TR03-296.
Dwivedi, A., Rehman, K., Ghosh, M. and Raman, R. (2018) Data Mining Algorithms in Healthcare. International Journal of Computer Applications, 180, 26-31. https://doi.org/10.5120/ijca2018916901
Guyon, I., Gunn, S., Nikravesh, M. and Zadeh, L.A. (2008) Feature Extraction: Foundations and Applications. Springer, Berlin.
Xue, B., Zhang, M. and Browne, W.N. (2013) Particle Swarm Optimization for Feature Selection in Classification: A Multi-Objective Approach. IEEE Transactions on Cybernetics, 43, 1656-1671. https://doi.org/10.1109/TSMCB.2012.2227469
Caldwell, J.C. and Caldwell, P. (2002) Africa: The New Family Planning Frontier. Studies in Family Planning, 33, 76-86. https://doi.org/10.1111/j.1728-4465.2002.00076.x
Fayyad, U. (1997) Data Mining and Knowledge Discovery in Databases: Implications for Scientific Databases. Proceedings. Ninth International Conference on Scientific and Statistical Database Management, Olympia, 11-13 August 1997, 2-11.
Chaurasia, A.R. (2014) Contraceptive Use in India: A Data Mining Approach. International Journal of Population Research, 2014, Article ID: 821436. https://doi.org/10.1155/2014/821436
Han, J. and Kamber, M. (2006) Data Mining Concepts and Techniques. Morgan Kaufmann Publishers, Burlington.
Berry, M.J. and Linoff, G. (1997) Data Mining Techniques: For Marketing, Sales and Customer Support. Wiley, New York.
Parr Rud, O. (2001) Data Mining Cookbook: Modeling Data for Marketing, Risk, and Customer Relationship Management. Wiley, New York.
Azevedo, A. and Santos, M.F. (2008) KDD, SEMMA and CRISP-DM: A Parallel Overview. IADIS European Conference on Data Mining 2008, Amsterdam, 24-26 July 2008, 182-185.
Suryani, D., Labellapansa, A. and Marsela, E. (2018) Accuracy of Algorithm C4.5 to Study Data Mining against Selection of Contraception. In: Saian, R. and Abbas, M., Eds., Proceedings of the Second International Conference on the Future of ASEAN (ICoFA) 2017—Volume 2, Springer, Singapore, 955-962. https://doi.org/10.1007/978-981-10-8471-3_95
Dwi Fajar Maulana, Y., Ruldeviyani, Y. and Indra Sensuse, D. (2020) Data Mining Classification Approach to Predict the Duration of Contraceptive Use. 2020 Fifth International Conference on Informatics and Computing (ICIC), Gorontalo, 3-4 November 2020, 1-6. https://doi.org/10.1109/ICIC50835.2020.9288568
Hailemariam, T., Gebregiorgis, A., Meshesha, M. and Mekonnen, W. (2017) Application of Data Mining to Predict the Likelihood of Contraceptive Method Use among Women Aged 15-49 Case of 2005 Demographic Health Survey Data Collected by Central Statistics Agency, Addis Ababa, Ethiopia. Journal of Health & Medical Informatics, 8, Article ID: 1000274. https://doi.org/10.4172/2157-7420.1000274
Witten, I.H., Frank, E. and Hall, M.A. (2011) Data Mining Practical Machine Learning Tools and Techniques. Morgan Kaufmann Publisher, Burlington.
Daelemans, W., Hoste, V., Meulder, F.D. and Naudts, B. (2003) Combined Optimization of Feature Selection and Algorithm Parameter Interaction in Machine Learning of Language. In: Lavrač, N., Gamberger, D., Blockeel, H. and Todorovski, L., Eds., Machine Learning: ECML 2003, Springer, Berlin, 84-95. https://doi.org/10.1007/978-3-540-39857-8_10
Chapman, P., Clinton, J., Kerber, R., Khabaza, T., Reinartz, T., Shearer, C. and Wirth, R. (2000) CRISP-DM 1.0: Step by Step Data Mining Guide. https://books.google.com/books/about/CRISP_DM_1_0.html?id=po7FtgAACAAJ
Liu, H. and Motoda, H. (1998) Feature Selection for Knowledge Discovery and Data Mining. Springer, New York. https://doi.org/10.1007/978-1-4615-5689-3
Brachman, R.J. and Anand, T. (1996) The Process of Knowledge Discovery in Databases. Advances in Knowledge Discovery and Data Mining, AAAI Press/the MIT Press, Menlo Park, 37-57. https://dl.acm.org/doi/10.5555/257938.257944
Yahia, M.E. and Ibrahim, B.A. (2003) K-Nearest Neighbor and C4.5 Algorithms as Data Mining Methods: Advantages and Difficulties. ACS/IEEE International Conference on Computer Systems and Applications, 2003. Book of Abstracts, Tunis, 14-18 July 2003, 103-109. https://doi.org/10.1109/AICCSA.2003.1227535
Bach, M.P. and Ćosić, D. (2008) Data Mining Usage in Health Care Management: Literature Survey and Decision Tree Application. Medicinski Glasnik, 5, 57-64.
Chakrabarti, S., Cox, E., Frank, E., Hartmut, G.R., Han, J., Jiang, X., Kamber, M. and Witten, I. (2009) Data Mining: Know It All. Morga Kaufmann Publishers, Burlington.
Famili, A. and Turney, P. (1997) Data Preprocessing and Intelligent Data Analysis. Institute of Information Technology, National Research Council Canada.
Lu, H., Setiono, R. and Liu, H. (1996) Effective Data Mining Using Neural Networks. IEEE Transactions on Knowledge and Data Engineering, 8, 957-961. https://doi.org/10.1109/69.553163
Witten, I.H. and Frank, E. (2005) Data Mining: Practical Machine Learning Tools and Techniques. 2nd Edition, Morgan Kaufmann Publishers, San Francisco.
Breiman, L., Friedman, J.H., Olshen, R.A. and Stone, C.J. (1984) Classification and Regresion Trees.
Velickov, S. and Solomatine, D. (2000) Predictive Data Mining: Practical Examples: Artificial Intelligence in Civil Engineering. 2nd Joint Workshop on Applied AI in Civil Engineering, Cottbus, Germany, Cottbus, March 2000, 1-17.