Optimization of Skin Disease Classification with a Hybrid Convolutional Neural Network and Vision Transformer Approach
- 1 Department of Computer Science, Ignatius Ajuru University of Education, Port Harcourt, Nigeria
- 2 Department of Computer Science, Ignatius Ajuru University of Education, Port Harcourt, Nigeria
Abstract
Skin diseases remain a significant global health concern, with accurate diagnosis often challenged by visual similarities among lesion types and limited access to specialist dermatological services. This study presents an optimized hybrid Convolutional Neural Network-Vision Transformer (CNN-ViT) framework for automated skin disease classification using dermoscopic images from the HAM10000 dataset. The proposed architecture combines the local feature extraction capability of CNNs with the global contextual learning ability of Vision Transformers to improve classification performance. Prior to model training, the dataset was partitioned into training (70%), validation (15%), and testing (15%) subsets, with data augmentation applied exclusively to the training set to address class imbalance and improve generalization. The hybrid model was implemented using TensorFlow/Keras and evaluated using weighted accuracy, precision, recall, and F1-score metrics. Experimental results demonstrated a classification accuracy of 97.68%, precision of 89.80%, recall of 90.20%, and F1-score of 91.30%. Comparative analysis indicated that the hybrid architecture outperformed conventional CNN-based approaches due to its ability to simultaneously capture local texture patterns and global lesion structures. The findings demonstrate the effectiveness of integrating CNN and ViT architectures for intelligent dermatological image analysis and highlight the potential of hybrid deep learning models in supporting clinical decision-making and early skin disease diagnosis.
- Del Rosso, J.Q. and Levin, J. (2011) The Clinical Relevance of Maintaining the Functional Integrity of the Stratum Corneum in Both Healthy and Disease-Affected Skin. Journal of Clinical and Aesthetic Dermatology , 4, 22-42.
- Apalla, Z., Nashan, D., Weller, R.B. and Castellsagué, X. (2017) Skin Cancer: Epidemiology, Disease Burden, Pathophysiology, Diagnosis, and Therapeutic Approaches. Dermatology and Therapy , 7, 5-19. https://doi.org/10.1007/s13555-016-0165-y
- Litjens, G., Kooi, T., Bejnordi, B.E., Setio, A.A.A., Ciompi, F., Ghafoorian, M., et al . (2017) A Survey on Deep Learning in Medical Image Analysis. Medical Image Analysis , 42, 60-88. https://doi.org/10.1016/j.media.2017.07.005
- Tschandl, P., Rosendahl, C. and Kittler, H. (2018) The HAM10000 Dataset, a Large Collection of Multi-Source Dermatoscopic Images of Common Pigmented Skin Lesions. Scientific Data , 5, Article No. 180161. https://doi.org/10.1038/sdata.2018.161
- Aydin, Y. (2023) A Comparative Analysis of Skin Cancer Detection Applications Using Histogram-Based Local Descriptors. Diagnostics , 13, Article 3142. https://doi.org/10.3390/diagnostics13193142
- Codella, N., Rotemberg, V., Tschandl, P., Celebi, M.E., Dusza, S., Gutman, D., et al . (2018) Skin Lesion Analysis toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC). arXiv: 1902.03368.
- Bengio, Y., Courville, A. and Vincent, P. (2013) Representation Learning: A Review and New Perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence , 35, 1798-1828. https://doi.org/10.1109/tpami.2013.50
- Brinker, T.J., Hekler, A., Enk, A.H., Berking, C., Haferkamp, S., Hauschild, A., et al . (2019) Deep Neural Networks Are Superior to Dermatologists in Melanoma Image Classification. European Journal of Cancer , 119, 11-17. https://doi.org/10.1016/j.ejca.2019.05.023
- LeCun, Y., Bengio, Y. and Hinton, G. (2015) Deep Learning. Nature , 521, 436-444. https://doi.org/10.1038/nature14539
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł. and Polosukhin, I. (2017) Attention Is All You Need. Proceedings of the 31 st Conference on Neural Information Processing Systems ( NeurIPS 2017), Long Beach, 4-9 December 2017, 5998-6008.
- Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al . (2021) An Image Is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations ( ICLR 2021), Vienna, 4 May 2021.