A Comparative Survey on Arabic Stemming: Approaches and Challenges
Department of Computer Information Systems, Faculty of Computers and Information Technology, University of Tabuk, Tabuk, SA
,
Department of Computer Science, College of Computer Science and Information Technology, Sudan University of Science and Technology, Khartoum State, Sudan
,
Department of Computer Information Systems, School of Information Technology, Al-Balqa Applied University, Salt, Jordan
,
Department of Information Technology, Faculty of Computers and Information Technology, University of Tabuk, Tabuk, Saudi Arabia
1 Department of Computer Information Systems, Faculty of Computers and Information Technology, University of Tabuk, Tabuk, SA
2 Department of Computer Science, College of Computer Science and Information Technology, Sudan University of Science and Technology, Khartoum State, Sudan
3 Department of Computer Information Systems, School of Information Technology, Al-Balqa Applied University, Salt, Jordan
4 Department of Information Technology, Faculty of Computers and Information Technology, University of Tabuk, Tabuk, Saudi Arabia
Arabic, as one of the Semitic languages, has a very rich and complex morphology, which is radically different from the European and the East Asian languages. The derivational system of Arabic, is therefore, based on roots, which are often inflected to compose words, using a spectacular and a relatively large set of Arabic morphemes affixes, e.g., antefixs, prefixes, suffixes, etc. Stemming is the process of rendering all the inflected forms of word into a common canonical form. Stemming is one of the early and major phases in natural processing, machine translation and information retrieval tasks. A number of Arabic language stemmers were proposed. Examples include light stemming, morphological analysis, statistical-based stemming, N-grams and parallel corpora (collections). Motivated by the reported results in the literature, this paper attempts to exhaustively review current achievements for stemming Arabic texts. A variety of algorithms are discussed. The main contribution of the paper is to provide better understanding among existing approaches with the hope of building an error-free and effective Arabic stemmer in the near future.
Manning, C.D., Raghavan, P. and Schutze, H. (2008) Introduction to Information Retrieval. Cambridge University Press, Cambridge. https://doi.org/10.1017/CBO9780511809071
Mustafa (2013) Mixed-Language Arabic-English Information Retrieval. PhD Thesis, University of Cape Town, Cape Town.
Darwish, K. and Magdy, W. (2014) Arabic Information Retrieval. Foundations and Trends in Information Retrieval, 7, 239-342. https://doi.org/10.1561/1500000031
Larkey, L., Ballesteros, L. and Connell, M. (2007) Light Stemming for Arabic Information Retrieval. In: Soudi, A., van den Bosch, A. and Neumann, G., Eds., Arabic Computational Morphology, Springer, Berlin, 221-243. https://doi.org/10.1007/978-1-4020-6046-5_12
Pirkola, A., Hedlund, T., Keskustalo, H. and Jarvelin, K. (2001) Dictionary-Based Cross-Language Information Retrieval: Problems, Methods, and Research Findings. Information Retrieval, 4, 209-230. https://doi.org/10.1023/A:1011994105352
Mirkin, B. (2010) Population Levels, Trends and Policies in the Arab Region: Challenges and Opportunities. Arab Human Development, Report Paper 1.
Cheung, W. (2008) Web Searching in a Multilingual World. Communications of the ACM, 51, 32-40. https://doi.org/10.1145/1342327.1342335
Habash, N. and Rambow, O. (2007) Arabic Diacritization through Full Morphological Tagging. Human Language Technologies: The Conference of the North American Chapter of the Association for Computational Linguistics, Rochester, 22-27 April 2007, 53-56.
Manzour, I. (2017) Lisan Al-Arab. www.lesanarab.com
Hegazi, N. and El-sharkawi, A. (1985) An Approach to a Computerized Lexical Analyzer for Natural Arabic Text. Proceedings of the Arabic Language Conference, Kuwait, 14-16 April 1985.
Mustafa, M. and Suleman, H. (2011) Building a Multilingual and Mixed Arabic-English Collection. 3rd Arabic Language Technology International Conference, Alexandria, 17-18 July 2011, 28-37.
Kadri, Y. and Nie, J.Y. (2006) Effective Stemming for Arabic Information Retrieval. Proceedings of the Challenge of Arabic for NLP/MT Conference, London, 23 October 2006, 68-74.
Attia, M.A. (2008) Handling Arabic Morphological and Syntactic Ambiguity within the LFG Framework with a View to Machine Translation. PhD Thesis, The University of Manchester, Manchester.
Attia, M.A. (2007) Arabic Tokenization System. Proceedings of the 2007 Workshop on Computational Approaches to Semitic Languages: Common Issues and Resources, Prague, 28 June 2007, 65-72. https://doi.org/10.3115/1654576.1654588
A Comparative Survey on Arabic Stemming: Approaches and Challenges — Oak Academic Publishing
Goweder, A., Poesio, M., De Roeck, A. and Reynolds, J. (2005) Identifying Broken Plurals in Unvowelised Arabic Text. Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing, Vancouver, 6-8 October 2005, 246-253.
Buckwalter, T. (2004) Issues in Arabic Orthography and Morphology Analysis. Proceedings of the Workshop on Computational Approaches to Arabic Script-Based Languages, Geneva, 28 August 2004, 31-34. https://doi.org/10.3115/1621804.1621813
Levow, G.A., Oard, D.W. and Resnik, P. (2005) Dictionary-Based Techniques for Cross-Language Information Retrieval. Information Processing & Management, 41, 523-547. https://doi.org/10.1016/j.ipm.2004.06.012
Abdelali, A. (2006) Improving Arabic Information Retrieval Using Local Variations in Modern Standard Arabic. PhD Thesis, New Mexico Institute of Mining and Technology, New Mexico.
Darwish, K. and Oard, D.W. (2003) CLIR Experiments at Maryland for TREC-2002: Evidence Combination for Arabic-English Retrieval.
Daoud, D. and Hasan, Q. (2011) Stemming Arabic Using Longest-Match and Dynamic Normalization. 3rd Arabic Language Technology International Conference, Alexandria, 17-18 July 2011, 54-59.
Xu, J., Fraser, A. and Weischedel, R. (2001) TREC 2001 Cross-Lingual Retrieval at BBN. Text Retrieval Conference, Gaithersburg, 13-16 November 2001, 68-78.
Xu, J., Fraser, A. and Weischedel, R. (2002) Empirical Studies in Strategies for Arabic Retrieval. Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Tampere, 11-15 August 2002, 269-274. https://doi.org/10.1145/564376.564424
Khoja, S. and Garside, R. (1999) Stemming Arabic Text. Computing Department, Lancaster University, Lancaster.
Hammo, B.H. (2009) Towards Enhancing Retrieval Effectiveness of Search Engines for Diacritisized Arabic Documents. Information Retrieval, 12, 300-323. https://doi.org/10.1007/s10791-008-9081-9
Darwish, K. (2002) Building a Shallow Arabic Morphological Analyzer in One Day. Proceedings of the ACL Workshop on Computational Approaches to Semitic Languages, Morristown, July 2002, 1-8. https://doi.org/10.3115/1118637.1118643
Buckwalter, T. (2002) Buckwalter Arabic Morphological Analyzer Version 1.0. Linguistic Data Consortium, University of Pennsylvania, Philadelphia.
Ghwanmeh, S., Kanaan, G., Al-Shalabi, R. and Alrababah, S. (2009) Enhanced Algorithm for Extracting the Root of Arabic Words. 6th International Conference on Computer Graphics, Imaging and Visualization, Tian Jin, 11-14 August 2009, 388-391.
Al-Kabi, M., Kazakzeh, S., Abu Ata, B., Al-Rababah, A.S. and Izzat, A.M. (2015) A Novel Root Based Arabic Stemmer. Journal of King Saud University—Computer and Information Sciences, 27, 94-103. https://doi.org/10.1016/j.jksuci.2014.04.001
Aljlayl, M. and Frieder, O. (2002) On Arabic Search: Improving the Retrieval Effectiveness via Light Stemming Approach. Proceedings of the 11th ACM International Conference on Information and Knowledge Management, Illinois, 4-9 November 2002, 340-347. https://doi.org/10.1145/584792.584848
Chen, A. and Gey, F. (2002) Building an Arabic Stemmer for Information Retrieval. Text Retrieval Conference, Gaithersburg, 19-22 November 2002, 631-639.
Larkey, L.S., Ballesteros, L. and Connell, M.E. (2002) Improving Stemming for Arabic Information Retrieval: Light Stemming and Co-Occurrence Analysis. In: Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Tampere, 11-15 August 2002, 275-282. https://doi.org/10.1145/564376.564425
Diab, M., Hacioglu, K. and Jurafsky, D. (2004) Automatic Tagging of Arabic Text: From Raw Text to Base Phrase Chunks. Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, Boston, 2-7 May 2004, 149-152. https://doi.org/10.3115/1613984.1614022
Nwesri, A.F.A., Tahaghoghi, S.M.M. and Scholer, F. (2005) Stemming Arabic Conjunctions and Prepositions. Lecture Notes in Computer Science, 3772, 206-217. https://doi.org/10.1007/11575832_23
Nwesri, A., Tahaghoghi, S.M.M. and Scholer, F. (2007) Arabic Text Processing for Indexing and Retrieval. Proceedings of the International Colloquium on Arabic Language Processing, Rabat, 18-19 June 2007, 18-19.
Ababneh, M., Al-Shalabi, R., Kanaan, G. and Al-Nobani, A. (2012) Building an Effective Rule-Based Light Stemmer for Arabic Language to Improve Search Effectiveness. International Arab Journal of Information Technology, 9, 368-372.
Sameer, R. (2016) Modified Light Stemming Algorithm for Arabic Language. Iraqi Journal of Science, 57, 507-513.
Habash, N. and Rambow, O. (2006) MAGEAD: A Morphological Analyzer and Generator for the Arabic Dialects. Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics, Sydney, 17-21 July 2006, 681-688. https://doi.org/10.3115/1220175.1220261
Salloum, W. and Habash, N. (2014) ADAM: Analyzer for Dialectal Arabic Morphology. Journal of King Saud University—Computer and Information Sciences, 26, 372-378. https://doi.org/10.1016/j.jksuci.2014.06.010
Abu Ata, B. and Al-Omari, A. (2014) A Rule-Based Stemmer for Arabic Gulf Dialect. Journal of King Saud University—Computer and Information Sciences, 27, 104-112.
Xu, J. and Croft, W.B. (1998) Corpus-Based Stemming Using Cooccurrence of Word Variants. ACM Transactions on Information Systems, 16, 61-81. https://doi.org/10.1145/267954.267957
Mustafa, S.H. and Al-Radaideh, Q.A. (2004) Using N-Grams for Arabic Text Searching. Journal of the American Society for Information Science and Technology, 55, 1002-1007. https://doi.org/10.1002/asi.20051
Khreisat, L. (2006) Arabic Text Classification Using N-Gram Frequency Statistics a Comparative Study. Proceedings of the 2006 International Conference on Data Mining, Las Vegas, 26-29 June 2006, 78-82.
Hmeidi, I.I., Al-Shalabi, R.F., Al-Taani, A.T., Najadat, H. and Al-Hazaimeh, S.A. (2010) A Novel Approach to the Extraction of Roots from Arabic Words Using Bigrams. Journal of the Association for Information Science and Technology, 61, 583-591.
Mansour, N., Haraty, R.A., Daher, W. and Houri, M. (2008) An Auto-Indexing Method for Arabic Text. Information Processing & Management, 44, 1538-1545. https://doi.org/10.1016/j.ipm.2007.12.007
Och, F.J. and Ney, H. (2003) A Systematic Comparison of Various Statistical Alignment Models. Computational Linguistics, 29, 19-51. https://doi.org/10.1162/089120103321337421
Al-shammari, E.T. and Lin, J. (2008) Towards an Error-Free Arabic Stemming. Proceedings of the 2nd ACM Workshop on Improving Non English Web Searching, Napa Valley, 26-30 October 2008, 9-16. https://doi.org/10.1145/1460027.1460030
Al-Serhan, H. and Ayesh, A. (2006) A Triliteral Word Roots Extraction Using Neural Network for Arabic. International Conference on Computer Engineering and Systems, Cairo, 5-7 November 2006, 436-440. https://doi.org/10.1109/icces.2006.320487
Boubas, A., Leena, L.L., Belkhouche, B. and Harous, S. (2011) GENESTEM: A Novel Approach for an Arabic Stemmer Using Genetic Algorithms. International Conference on Innovations in Information Technology, Abu Dhabi, 25-27 April 2011, 77-82. https://doi.org/10.1109/innovations.2011.5893872