A Multi-Modal Approach for Arabic Sign Language Gesture Recognition Using Deep Learning
- 1 College of Computer Science and Engineering, Taibah University, Madinah, Saudi Arabia
Abstract
This paper proposes a multi-modal deep learning framework for Arabic Sign Language (ArSL) recognition, addressing the challenges of both static and dynamic gesture recognition. The framework integrates spatial, temporal, and depth features using CNN, Transformer, and Depth-CNN models, combined via an attention-based fusion mechanism. A hierarchical recognition approach first classifies gestures as static or dynamic, then processes them with specialized models: MobileNetV3 for dynamic gestures and an MLP-KAN hybrid for static gestures. Evaluated on four ArSL datasets (Kaggle ASL, ArSL2018, DArSL50, KSU-ArSL), the system achieves 98.4% overall accuracy with real-time inference speeds of 0.007 seconds for static gestures and 0.02 seconds for dynamic gestures. Ablation studies confirm the importance of multi-modal fusion, with attention-based fusion improving accuracy by 11% compared to simple concatenation. The system demonstrates strong generalization across diverse datasets and conditions, making it suitable for real-world deployment in assistive communication technologies.
- World Health Organization (2025) Deafness and Hearing Loss. https://www.who.int/news-room/fact-sheets/detail/deafness-and-hearing-loss
- Center for Strategic and International Studies (2025) Disability Inclusion in Foreign Policy: Special Advisor Sara Minkara. https://www.csis.org/events/disability-inclusion-foreign-policy-special-advisor-sara-minkara
- Shin, J., Miah, A.S.M., Kabir, M.H., Rahim, M.A. and Al Shiam, A. (2024) A Methodological and Structural Review of Hand Gesture Recognition across Diverse Data Modalities. IEEE Access , 12, 142606-142639. https://doi.org/10.1109/access.2024.3456436
- Al Abdullah, B.A., Amoudi, G.A. and Alghamdi, H.S. (2024) Advancements in Sign Language Recognition: A Comprehensive Review and Future Prospects. IEEE Access , 12, 128871-128895. https://doi.org/10.1109/access.2024.3457692
- Noor, T.H., Noor, A., Alharbi, A.F., Faisal, A., Alrashidi, R., Alsaedi, A.S., et al. (2024) Real-Time Arabic Sign Language Recognition Using a Hybrid Deep Learning Model. Sensors , 24, Article 3683. https://doi.org/10.3390/s24113683
- Duwairi, R.M. and Halloush, Z.A. (2022) Automatic Recognition of Arabic Alphabets Sign Language Using Deep Learning. International Journal of Electrical and Computer Engineering ( IJECE ), 12, 2996-3004. https://doi.org/10.11591/ijece.v12i3.pp2996-3004
- Alharthi, N.M. and Alzahrani, S.M. (2023) Vision Transformers and Transfer Learning Approaches for Arabic Sign Language Recognition. Applied Sciences , 13, Article 11625. https://doi.org/10.3390/app132111625
- Alsolai, H., Alsolai, L., Al-Wesabi, F.N., Othman, M., Rizwanullah, M. and Abdelmageed, A.A. (2024) Automated Sign Language Detection and Classification Using Reptile Search Algorithm with Hybrid Deep Learning. Heliyon , 10, e23252. https://doi.org/10.1016/j.heliyon.2023.e23252
- Kong, F., Hu, K., Li, Y., Li, D., Liu, X. and Durrani, T.S. (2022) A Spectral-Spatial Feature Extraction Method with Polydirectional CNN for Multispectral Image Compression. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , 15, 2745-2758. https://doi.org/10.1109/jstars.2022.3158281
- Abdul Ameer, R.S., Ahmed, M.A., Al-Qaysi, Z.T., Salih, M.M. and Shuwandy, M.L. (2024) Empowering Communication: A Deep Learning Framework for Arabic Sign Language Recognition with an Attention Mechanism. Computers , 13, Article 153. https://doi.org/10.3390/computers13060153
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, L. and Polosukhin, I. (2017) Attention Is All You Need. arXiv: 1706.03762.