The aim of this research is to develop a speech synthesis model tailored towards Nigerian languages by leveraging natural language processing tool such as FastSpeech 2 and meta-tts for high-quality, non-autoregressive text-to-speech (TTS) generation and HiFi-GAN for neural vocoding. It was motivated due to lack of high-quality synthetic speech models for low-resource languages especially Nigerian Languages and specificially Hausa, Igbo and Yoruba. The methodology adopted is a structured and iterative approach that integrates Structured System Analysis and Design Methodology (SSADM), and Machine Learning Development Lifecycle (MLDLC), which incorporates a feasibility study, corpus collection, phonetic analysis, and model training with Nigerian Language annotated speech dataset. Speech dataset will be collected and preprocess from selected Nigerian languages(Igbo Hausa and Yoruba). Phonemes will be developed using both rule-based approach and grapheme-to-phoneme models like Epitran and Phonemizer, while FastSpeech 2, meta-tts and HiFi-GAN will be fine-tuned to accommodate tonal variations and prosodic patterns inherent in these languages. The model training pipeline will integrate Tacotron-based aligners for efficient text-to-mel-spectrogram conversion, while HiFi-GAN will enhance naturalness and intelligibility. Python will be the primary programming language for the implementation of this research while the interface will be a combination of hypertext markup language (HTML), cascading style sheet (CSS) and Javascript. The expected outcome is a state-of-the-art, speech synthesis system capable of generating natural and intelligible speech across multiple Nigerian languages.
KeywordsLanguagesNatural Language ProcessingHiFi-GANModelsFastSpeech2Speech SynthesisMeta-TTSDeep Learning
Akujobi, O.S. (2019) The English Language Coalescence and Multilingualism in Nigeria. IGWEBUIKE : An African Journal of Arts and Humanities , 5, 1-16.
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z. and Liu, T.-Y. (2020) FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. http://arxiv.org/abs/2006.04558
Afolabi, A., Omidiora, E. and Arulogun, T. (2013) Development of Text to Speech System for Yoruba Language. Innovative Systems Design and Engineering , 4, 1-7. https://www.iiste.org
Ekpenyong, M.E., Urua, E. and Gibbon, D. (2008) Towards an Unrestricted Domain TTS System for African Tone Languages. International Journal of Speech Technology , 11, 87-96. https://doi.org/10.1007/s10772-009-9037-5
Oyelade, J.O., Isewon, I., Famade, A. and Oyelade, J. (2022) Foundation of Computer Science FCS. International Journal of Applied Information Systems , 12.
Kong, J., Kim, J. and Bae, J. (2020) HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. http://arxiv.org/abs/2010.05646
Ping, W., Peng, K., Zhao, K. and Song, Z. (2019) WaveFlow: A Compact Flow-Based Model for Raw Audio. http://arxiv.org/abs/1912.01219
Ngor, C.I.-A. (2024) Tone Nature of Nigerian English. African Journal of Humanities and Contemporary Education Research , 15, 399-415. https://doi.org/10.62154/e2bnwx92
Salau, A.O., Olowoyo, T.D. and Akinola, S.O. (2020) Accent Classification of the Three Major Nigerian Indigenous Languages Using 1D CNN LSTM Network Model. In: Jain, S., et al ., Eds., Advances in Computational Intelligence Techniques , Springer, 1-16. https://doi.org/10.1007/978-981-15-2620-6_1
Fromont, R., Clark, L., Black, J.W. and Blackwood, M. (2023) Maximizing Accuracy of Forced Alignment for Spontaneous Child Speech.
Wu, H., Yun, J., Li, X., Huang, H. and Liu, C. (2023) Using a Forced Aligner for Prosody Research. Humanities and Social Sciences Communications , 10, Article No. 429. https://doi.org/10.1057/s41599-023-01931-4
UNESCO (2021) Towards Sustainable Preservation and Accessibility of Documentary Heritage. https://unesdoc.unesco.org/ark:/48223/pf0000380171
Tan, Y. and Jehom, W.J. (2024) Preservation , Digital Technology & Culture , 53, 165-177. https://doi.org/10.1515/pdtc-2024-0021
Hunt, A.J. and Black, A.W. (1996) Unit Selection in a Concatenative Speech Synthesis System Using a Large Speech Database. 1996 IEEE International Conference on Acoustics , Speech , and Signal Processing Conference Proceedings , Vol. 1, 373-376. https://doi.org/10.1109/icassp.1996.541110
Rabiner, L.R. and Schafer, R.W. (2007) Introduction to Digital Speech Processing. Foundations and Trends® in Signal Processing , 1, 1-194. https://doi.org/10.1561/2000000001
Bäckström, T., Räsänen, O., Zewoudie, A., Zarazaga, P.P., Koivusalo, L., Das, S., et al . (2022) Introduction to Speech Processing: 2nd Edition. https://doi.org/10.5281/ZENODO.6821775
Kuligowska, K., Kisielewicz, P. and Włodarz, A. (2018) Speech Synthesis Systems: Disadvantages and Limitations. International Journal of Engineering & Technology , 7, 234-239. https://doi.org/10.14419/ijet.v7i2.28.12933
Hande, S.S. (2014) A Review of Concatenative Text to Speech Synthesis. International Journal of Latest Technology in Engineering , Management & Applied Science , 3, 12-15. https://www.academia.edu/download/34937201/12-15.pdf
Khan, R.A., and Chitode, J.S. (2016) Concatenative Speech Synthesis: A Review. International Journal of Computer Applications , 136, 1-6. https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&doi=dcb1aefcc8d80c90392fa9b6f2740b4516e8ec44
Wang, Y., Skerry-Ryan, R.J., Stanton, D., Wu, Y., Weiss, R.J., Jaitly, N., et al . (2017) Tacotron: Towards End-to-End Speech Synthesis. Interspeech 2017, Stockholm, 20-24 August 2017, 4006-4010. https://doi.org/10.21437/interspeech.2017-1452
Jia, Y., Zhang, Y., Weiss, R.J., Wang, Q., Shen, J., Ren, F., et al . (2018) Transfer Learning from Speaker Verification to Multispeaker Text-to-Speech Synthesis. http://arxiv.org/abs/1806.04558
Ohala, J.J. and Kawasaki, H. (1984) Prosodic Phonology and Phonetics. Phonology Yearbook , 1, 113-127. https://doi.org/10.1017/s0952675700000312
Peterson, G.E. and Shoup, J.E. (1966) A Physiological Theory of Phonetics. Journal of Speech and Hearing Research , 9, 5-67. https://doi.org/10.1044/jshr.0901.05
Pierrehumbert, J.B. (1980) The Phonology and Phonetics of English Intonation.
Fant, G. (1981) The Source Filter Concept in Voice Production. STL-QPSR, 22, 021-037. https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&doi=647bee8e1ea9b5fcaa27dd8c0937a165a8f5f717
Ma, R., Qian, M., Fathullah, Y., Tang, S., Gales, M. and Knill, K. (2025) Cross-Lingual Transfer Learning for Speech Translation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics : Human Language Technologies , Volume 2, 33-43. https://doi.org/10.18653/v1/2025.naacl-short.4
Fant, G. (2001) T. Chiba and M. Kajiyama, Pioneers in Speech Acoustics. Journal of the Phonetic Society of Japan , 5, 4-5. https://doi.org/10.24467/onseikenkyu.5.2_4
Yoshimura, T., Tokuda, K., Masuko, T., Kobayashi, T. and Kitamura, T. (1999) Simultaneous Modeling of Spectrum, Pitch and Duration in Hmm-Based Speech Synthesis. 6 th European Conference on Speech Communication and Technology , EUROSPEECH 1999, Budapest, 5-9 September 1999, 2347-2350.
Ping, W., Peng, K. and Chen, J. (2018) ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech. http://arxiv.org/abs/1807.07281
Atoi, N.E. (2024) Language and Communication Implication of Artificial Intelligence on Selected Nigerian University Undergraduates. UJAH : Unizik Journal of Arts and Humanities , 25, 109-154. https://doi.org/10.4314/ujah.v25i1.5
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L. and Polosukhin, I. (2017) Attention Is All You Need. Proceedings of the 31 st International Conference on Neural Information Processing Systems , Long Beach, 4-9 December 2017, 6000-6010. http://arxiv.org/abs/1706.03762
Bleyan, H., Ritchie, S., Mortensen, J.F. and Esch, D.V. (2019) Developing Pronunciation Models in New Languages Faster by Exploiting Common Grapheme-to-Phoneme Correspondences across Languages. Interspeech 2019, Graz, 15-19 September 2019, 2100-2104. https://doi.org/10.21437/interspeech.2019-1781
Hassana, I.L. and Sanusi, M. (2019) Text to Speech Synthesis System in Yoruba Language. International Journal of Advances in Scientific Research and Engineering , 5, 180-191. https://doi.org/10.31695/ijasre.2019.33568
Olaniyan, O.M. and Akinode, V. (2023) Development of a Text-to-Speech Synthesis for Yoruba Language Using Deep Learning. Technology and Innovation , 2, 1-7.
Shen, J., Pang, R., Weiss, R.J., Schuster, M., Jaitly, N., Yang, Z., et al . (2017) Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions. http://arxiv.org/abs/1712.05884
Artetxe, M., Ruder, S. and Yogatama, D. (2020) On the Cross-Lingual Transferability of Monolingual Representations. Proceedings of the 58 th Annual Meeting of the Association for Computational Linguistics , July 2020, 4623-4637. https://doi.org/10.18653/v1/2020.acl-main.421
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., et al . (2020) Unsupervised Cross-Lingual Representation Learning at Scale. Proceedings of the 58 th Annual Meeting of the Association for Computational Linguistics , July 2020, 8440-8451. https://doi.org/10.18653/v1/2020.acl-main.747
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., et al . (2016) WaveNet: A Generative Model for Raw Audio. http://arxiv.org/abs/1609.03499