Stemming is used to produce stem or root of words. The process is vital to different research fields such as text mining, sentiment analysis, and text categorization, etc. Several techniques have been proposed to stemming Arabic text and among them, Khoja and light-10 stemmers are the most widely used. In this paper, we propose and evaluate two different stemming techniques to Arabic that are based on light stemming techniques. The new stemmers are compared to best reported light stemmer, which is light-10. Results and experiments, which were conducted using standard collections, reveal that The proposed stemmers yield 5.13% and 13.1% improvement in retrieval performance over light 10 with 0.369 average precision and 0.397, respectively and the improvement is statistically significant.
KeywordsArabic LanguageArabic Information RetrievalLight StemmingLight 10Extended Light-StemmerLinguistic-Based Stemmer
Miniwatts Marketing Group (2018) Internet World Stats Usage and Population Statistics. http://www.internetworldstats.com/stats7.htm
Simpson, A.K. (2009) The Origin and Development of Nonconcatenative Morphology. Ph.D. Thesis, University of California, Berkeley.
Mustafa, M., Eldeen, A.S., Bani-Ahmad, S. and Elfaki, A.O. (2017) A Comparative Survey on Arabic Stemming: Approaches and Challenges. Intelligent Information Management, 2017, 39-67.
Saad, M.K. and Ashour, W. (2010) OSAC: Open Source Arabic Corpora. The 6th International Conference on Electrical and Computer Systems (EECS’10), Lefke, 25-26 November 2010, ,118-123,
Mustafa, M. (2013) Mixed-Language Arabic-English Information Retrieval. Ph.D. Thesis. University of Cape Town, Cape Town.
Manzour, I. (2018) Lisan Al-Arab. http://www.lesanarab.com/
Hegazi, N. and El-sharkawi, A. (1985) An Approach to a Computerized Lexical Analyzer for Natural Arabic Text. Proceedings of the Arabic Language Conference, Kuwait, 14-16.
Mustafa, M. and Suleman, H. (2011) Building a Multilingual and Mixed Arabic-English Collection. The Proceedings of the 3rd Arabic Language Technology International Conference (ALTIC), Alexandria, 9-11 October 2001.
Khoja, S. and Garside, R. (1999) Stemming Arabic Text. Computing Department, Lancaster University, Lancaster.
Buckwalter, T. (2002) Buckwalter Arabic Morphological Analyzer Version 1.0. Linguistic Data Consortium, University of Pennsylvania, Philadelphia.
Darwish, K. (2002) Building a Shallow Arabic Morphological Analyzer in One Day. Proceedings of the ACL Workshop on Computational Approaches to Semitic Languages, Philadelphia, 11 July 2002, 1-8. https://doi.org/10.3115/1118637.1118643
Xu, J., Fraser, A. and Weischedel, R. (2001) TREC 2001 Cross-Lingual Retrieval at BBN. TREC 2001, Gaithersburg, 13 November 2011, 68-78.
Al-Sughaiyer, I.A. and Al-Kharashi, I.A. (2006) Rule Parser for Arabic Stemmer. In: Sojka, P., Kopecek, I. and Pala, K., Eds., Text, Speech and Dialogue, TSD 2002. Lecture Notes in Computer Science, Vol. 2448, Springer, Berlin, Heidelberg. https://doi.org/10.1007/3-540-46154-X_2
Al-Shalabi, R., Kanaan, G., Ghwanmeh, S. and Nour, F.M. (2007) Stemmer Algorithm for Arabic Words Based on Excessive Letter Locations. 4th International Conference on Innovations in Information Technology (IIT ‘07), Dubai, 18-20 November 2007, 456-460. https://doi.org/10.1109/IIT.2007.4430444
Darwish, K. and Oard, D.W. (2003) CLIR Experiments at Maryland for TREC-2002: Evidence Combination for Arabic-English Retrieval. TREC 2003 Proceedings, College Park, February, 2003.
Aljlayl, M. and Frieder, O. (2002) On Arabic Search: Improving the Retrieval Effectiveness via Light Stemming Approach. In: Proceedings of the 11th ACM International Conference on Information and Knowledge Management, ACM Press, New York, 340-347. https://doi.org/10.1145/584792.584848
Kadri, Y. and Nie, J.Y. (2006) Effective Stemming for Arabic Information Retrieval. Proceedings of the Challenge of Arabic for NLP/MT Conference, Londres, 3 October 2006, 68-74.
Chen, A. and Gey, F. (2002) Building an Arabic Stemmer for Information Retrieval. In: TREC, NIST, Gaithersburg, 631-639.
Nwesri, A.F.A., Tahaghoghi, S.M.M. and Scholer, F. (2005) Stemming Arabic Conjunctions and Prepositions. Lecture Notes in Computer Science, 3772, 206-217. https://doi.org/10.1007/11575832_23
Larkey, L., Ballesteros, L. and Connell, M. (2007) Light Stemming for Arabic Information Retrieval. In: Arabic Computational Morphology, Springer, Berlin, 221-243.
Diab, M., Hacioglu, K. and Jurafsky, D. (2004) Automatic Tagging of Arabic Text: From Raw Text to Base Phrase Chunks. In: Proceedings of HLT-NAACL: Short Papers, Association for Computational Linguistics, 149-152. https://doi.org/10.3115/1613984.1614022
Ababneh, M., Al-Shalabi, R., Kanaan, G. and Al-Nobani, A. (2012) Building an Effective Rule-Based Light Stemmer for Arabic Language to Improve Search Effectiveness. The International Arab Journal of Information Technology, 9, 368-372.
Larkey, L.S., Ballesteros, L. and Connell, M.E. (2002) Improving Stemming for Arabic Information Retrieval: Light Stemming and Co-Occurrence Analysis. Annual ACM Conference on Research and Development in Information Retrieval: Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Tampere, 11-15 August 2002, 275-282. https://doi.org/10.1145/564376.564425
Xu, J. and Croft, W.B. (1998) Corpus-Based Stemming Using Co-Occurrence of Word Variants. ACM Transactions on Information Systems, 16, 61-81. https://doi.org/10.1145/267954.267957
Mustafa, S.H. and Al-Radaideh, Q.A. (2004) Using N-Grams for Arabic Text Searching. Journal of the American Society for Information Science and Technology, 55, 1002-1007. https://doi.org/10.1002/asi.20051
Xu, J., Fraser, A. and Weischedel, R. (2002) Empirical Studies in Strategies for Arabic Retrieval. Annual ACM Conference on Research and Development in Information Retrieval: Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Tampere, 11-15 August 2002, 269-274. https://doi.org/10.1145/564376.564424
Hmeidi, I.I., Al-Shalabi, R.F., Al-Taani, A.T., Najadat, H. and Al-Hazaimeh, S.A. (2010) A Novel Approach to the Extraction of Roots from Arabic Words Using Bigrams. Journal of American Society for Information Science and Technology, 61, 583-591.
Al-shammari, E.T. and Lin, J. (2008) Towards an Error-Free Arabic Stemming. In: Proceedings of the 2nd ACM workshop on Improving Non English Web Searching, ACM, New York, 9-16. https://doi.org/10.1145/1460027.1460030
Mansour, N., Haraty, R.A., Daher, W. and Houri, M. (2008) An Auto-Indexing Method for Arabic Text. Information Processing and Management, 44, 1538-1545. https://doi.org/10.1016/j.ipm.2007.12.007
Boubas, A., Lulu, L., Belkhouche, B. and Harous, S. (2011) GENESTEM: A Novel Approach for an Arabic Stemmer Using Genetic Algorithms. International Conference on Innovations in Information Technology, Abu Dhabi, 25-27 April 2011, 77-82. https://doi.org/10.1109/INNOVATIONS.2011.5893872
Al-Serhan, H. and Ayesh, A. (2006) A Triliteral Word Roots Extraction Using Neural Network for Arabic. International Conference on Computer Engineering and Systems, Cairo, 5-7 November 2006, 436-440. https://doi.org/10.1109/ICCES.2006.320487
Eldesouki, M., Arafa, A. and Darwish, K. (2009) Stemming Techniques of Arabic Language: Comparative Study from the Information Retrieval Perspective. The Egyptian Computer Journal, 36, 30-49.
The Stanford Natural Language Processing Group (2008) Arabic Natural Language Processing. https://nlp.stanford.edu/projects/arabic.shtml
Levow, G.A., Oard, D.W. and Resnik, P. (2005) Dictionary-Based Techniques for Cross-Language Information Retrieval. Information Processing and Management, 41, 523-547. https://doi.org/10.1016/j.ipm.2004.06.012