Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks
- 1 Applied College, Taibah University, Madinah, Saudi Arabia
Abstract
Machine learning models, and deep neural networks in particular, have achieved remarkable performance across a wide spectrum of tasks from image classification and natural language processing to autonomous navigation and medical diagnostics. However, alongside their successes, these models have exhibited a deep vulnerability: their susceptibility to adversarial examples, carefully crafted inputs designed to deceive a model into producing incorrect outputs. This paper presents a comprehensive survey and taxonomy of adversarial machine learning attacks and defenses, examining both the theoretical underpinnings and practical implications for real-world systems. We systematically classify attacks along four primary dimensions: knowledge of the target model (white-box versus black-box), the attack stage (training-time or inference-time), the attacker’s goal (targeted versus untargeted), and the perturbation modality (pixel-level, semantic, or physical). For each category, we review representative algorithms, analyze their strengths and limitations, and discuss the threat models they inhabit. On the defense side, we evaluate adversarial training, certified robustness methods, detection-based approaches, and input preprocessing techniques. We conduct a critical assessment of the arms race dynamic between attackers and defenders and propose a unified evaluation framework for comparing defense mechanisms across attack types. We also discuss open research challenges, including the transferability problem, robustness-accuracy trade-offs, and adversarial robustness in the physical world. Our goal is to provide a rigorous, accessible foundation for researchers and practitioners working at the intersection of machine learning, security, and systems design.
- Kebande, V.R., Ikuesan, R.A., Karie, N.M., Alawadi, S., Choo, K.R. and Al-Dhaqm, A. (2020) Quantifying the Need for Supervised Machine Learning in Conducting Live Forensic Analysis of Emergent Configurations (ECO) in IoT Environments. Forensic Science International : Reports , 2, Article 100122. https://doi.org/10.1016/j.fsir.2020.100122
- Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. and Fergus, R. (2014) Intriguing Properties of Neural Networks. arXiv: 1312.6199.
- Goodfellow, I., Shlens, J. and Szegedy, C. (2015) Explaining and Harnessing Adversarial Examples. arXiv: 1412.6572.
- Carlini, N. and Wagner, D. (2017) Towards Evaluating the Robustness of Neural Networks. 2017 IEEE Symposium on Security and Privacy ( SP ), San Jose, 22-26 May 2017, 39-57. https://doi.org/10.1109/sp.2017.49
- Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., et al . (2018) Robust Physical-World Attacks on Deep Learning Visual Classification. 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition , Salt Lake City, 18-23 June 2018, 1625-1634. https://doi.org/10.1109/cvpr.2018.00175
- Sharif, M., Bhagavatula, S., Bauer, L. and Reiter, M.K. (2016) Accessorize to a Crime. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security , Vienna, 24-28 October 2016, 1528-1540. https://doi.org/10.1145/2976749.2978392
- Xie, C., Wang, J., Zhang, Z., Zhou, Y., Xie, L. and Yuille, A. (2017) Adversarial Examples for Semantic Segmentation and Object Detection. 2017 IEEE International Conf erence on Computer Vision ( ICCV ), Venice, 22-29 October 2017, 1378-1387. https://doi.org/10.1109/iccv.2017.153
- Cohen, J., Rosenfeld, E. and Kolter, J. Z. (2019) Certified Adversarial Robustness via Randomized Smoothing. arXiv: 1902.02918.
- Mirman, M., Gehr, T. and Vechev, M. (2018) Differentiable Abstract Interpretation for Provably Robust Neural Networks. Proceedings of the 35 th International Conference on Machine Learning , Stockholm, 2018, 3578-3586.
- Madry, A., Makelov, A., Schmidt, L., Tsipras, D. and Vladu, A. (2018) Towards Deep Learning Models Resistant to Adversarial Attacks. arXiv: 1706.06083.
- Croce, F. and Hein, M. (2020) Reliable Evaluation of Adversarial Robustness with an Ensemble of Diverse Parameter-Free Attacks. International Conference on Machine Learning ( ICML ), Online, 13-18 July 2020 2206-2216.