Research ArticleOpen AccessGoogle Scholar indexed
Comparative Analysis of Neural Networks and Naive Bayes for Multilingual Text Identification
Department of Computer Science, Rochester Institute of Technology, Rochester, USA
Department of Computer Science, Rutgers University, New Brunswick, USA
Department of Computer Science, Rochester Institute of Technology, Rochester, USA
- 1 Department of Computer Science, Rochester Institute of Technology, Rochester, USA
- 2 Department of Computer Science, Rutgers University, New Brunswick, USA
- 3 Department of Computer Science, Rochester Institute of Technology, Rochester, USA
Journal of Software Engineering and Applications·Volume 18 (2025)·Pages 446–457·Published 18 November 2025·DOI10.4236/jsea.2025.1811027
Copy link · social · email
Abstract
This study presents a comparative analysis of two distinct machine learning approaches for multilingual text identification: character-level neural networks (CNN/RNN) and traditional Naive Bayes classifiers. We constructed a dataset comprising 20 languages, including Arabic, Bulgarian, German, Greek, English, Spanish, French, Hindi, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Russian, Swahili, Thai, Turkish, Urdu, and Vietnamese. Experimental results demonstrate that the character-level neural network model achieved 98.76.
KeywordsLanguage IdentificationCharacter-Level Neural NetworksNaive BayesMultilingual Text ClassificationCNN/RNNCharacter N-GramsComputational EfficiencyPerformance AnalysisText ProcessingMachine LearningNatural Language ProcessingComparative Study
- Cavnar, W.B. and Trenkle, J.M. (1994) N-Gram-Based Text Categorization. Proceedings of SDAIR -94, 3 rd Annual Symposium on Document Analysis and Information Retrieval , Las Vegas, 11-13 April 1994, 161-175.
- Dunning, T. (1994) Statistical Identification of Language. Computing Research Laboratory Technical Report, New Mexico State University.
- Ljubesic, N., Mikelic, N. and Boras, D. (2007) Language Indentification: How to Distinguish Similar Languages? 2007 29 th International Conference on Information Technology Interfaces , Cavtat, 25-28 June 2007, 541-546. https://doi.org/10.1109/iti.2007.4283829
- King, B. and Abney, S. (2013) Labeling the Languages of Words in Mixed-Language Documents Using Weakly Supervised Methods. Proceedings of NAACL - HLT 2013, Atlanta, Georgia, USA, 9-14 June 2013, 1110-1119.
- Jauhiainen, T., Linden, K. and Jauhiainen, H. (2017) Evaluation of Language Identification Methods Using 285 Languages. Proceedings of the 21 st Nordic Conference on Computational Linguistics ( NoDaLiDa ), Gothenburg, 22-24 May 2017, 183-191.
- Zhang, X., Zhao, J. and LeCun, Y. (2015) Character-Level Convolutional Networks for Text Classification. Advances in Neural Information Processing Systems ( NeurIPS 2015), Montréal, Quebec, Canada, 7-12 December 2015, 649-657.
- Kim, Y., Jernite, Y., Sontag, D. and Rush, A. (2016) Character-Aware Neural Language Models. Proceedings of the AAAI Conference on Artificial Intelligence , 30, 2741-2749. https://doi.org/10.1609/aaai.v30i1.10362
- Kocmi, T. and Bojar, O. (2017) LanideNN: Multilingual Language Identification on Character Window. Proceedings of the 14 th International Conference on Natural Language Processing ( ICON 2017), Kolkata, 18-21 December 2017, 194-201.
- Ali, A., Mubarak, H., and Durrani, N. (2021) A Comparative Study of Character-Level CNN and RNN Models for Language Identification. IEEE Access , 9, 58702-58713.
- Joulin, A., Grave, E., Bojanowski, P. and Mikolov, T. (2017) Bag of Tricks for Efficient Text Classification. Proceedings of the 15 th Conference of the European Chapter of the Association for Computational Linguistics : Volume 2, Short Papers , Valencia, 3-7 April 2017, 427-431. https://doi.org/10.18653/v1/e17-2068
- Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., et al. (2020) Unsupervised Cross-Lingual Representation Learning at Scale. Proceedings of the 58 th Annual Meeting of the Association for Computational Linguistics , 6-8 July 2020, 8440-8451. https://doi.org/10.18653/v1/2020.acl-main.747