Robustness, Cost, and Attack-Surface Concentration in Phishing Detection
- 1 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 2 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 3 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 4 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 5 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 6 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
- 7 Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, NC, USA
Abstract
Phishing detectors built on engineered website features attain near-perfect accuracy under i.i.d. evaluation, yet deployment security depends on robustness to post-deployment feature manipulation. We study this gap through a cost-aware evasion framework that models discrete, monotone feature edits under explicit attacker budgets. Three diagnostics are introduced: minimal evasion cost (MEC), the evasion survival rate S ( B ) , and the robustness concentration index (RCI). On the UCI Phishing Websites benchmark (11,055 instances, 30 ternary features), Logistic Regression, Random Forests, Gradient Boosted Trees, and XGBoost all achieve AUC ≥ 0.979 under static evaluation. Under budgeted sanitization-style evasion, robustness converges across architectures: the median MEC equals 2 with full features, and over 80% of successful minimal-cost evasions concentrate on three low-cost surface features. Feature restriction improves robustness only when it removes all dominant low-cost transitions. Under strict cost schedules, infrastructure-leaning feature sets exhibit 17% - 19% infeasible mass for ensemble models, while the median MEC among evadable instances remains unchanged. We formalize this convergence: if a positive fraction of correctly detected phishing instances admit evasion through a single feature transition of minimal cost c min , no classifier can raise the corresponding MEC quantile above c min without modifying the feature representation or cost model. Adversarial robustness in phishing detection is governed by feature economics rather than model complexity.
- Basit, A., Zafar, M., Liu, X., Javed, A.R., Jalil, Z. and Kifayat, K. (2020) A Comprehensive Survey of AI-Enabled Phishing Attacks Detection Techniques. Telecommunica tion Systems , 76, 139-154. https://doi.org/10.1007/s11235-020-00733-2
- Mohammad, R.M., Thabtah, F. and McCluskey, L. (2014) Predicting Phishing Websites Based on Self-Structuring Neural Network. Neural Computing and Applications , 25, 443-458. https://doi.org/10.1007/s00521-013-1490-z
- Sahingoz, O.K., Buber, E., Demir, O. and Diri, B. (2019) Machine Learning Based Phishing Detection from URLs. Expert Systems with Applications , 117, 345-357. https://doi.org/10.1016/j.eswa.2018.09.029
- Do, N.Q., Selamat, A., Krejcar, O., Herrera-Viedma, E. and Fujita, H. (2022) Deep Learning for Phishing Detection: Taxonomy, Current Challenges and Future Directions. IEEE Access , 10, 36429-36463. https://doi.org/10.1109/access.2022.3151903
- Biggio, B. and Roli, F. (2018) Wild Patterns: Ten Years after the Rise of Adversarial Machine Learning. Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security , Toronto, 15-19 October 2018, 2154-2156. https://doi.org/10.1145/3243734.3264418
- Apruzzese, G., Anderson, H.S., Dambra, S., Freeman, D., Pierazzi, F. and Roundy, K. (2023) “Real Attackers Don’t Compute Gradients”: Bridging the Gap between Adversarial ML Research and Practice. 2023 IEEE Conference on Secure and Trustworthy Machine Learning ( SaTML ), Raleigh, 8-10 February 2023, 339-364. https://doi.org/10.1109/satml54575.2023.00031
- Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., et al . (2013) Evasion Attacks against Machine Learning at Test Time. In: Lecture Notes in Computer Science , Springer, 387-402. https://doi.org/10.1007/978-3-642-40994-3_25
- Khonji, M., Iraqi, Y. and Jones, A. (2013) Phishing Detection: A Literature Survey. IEEE Communications Surveys & Tutorials , 15, 2091-2121. https://doi.org/10.1109/surv.2013.032213.00009
- Lin, Y., Liu, R., Divakaran, D.M., Ng, J.Y., et al . (2021) Phishpedia: A Hybrid Deep Learning Based Approach to Visually Identify Phishing Webpages. Proceedings of the 30 th USENIX Security Symposium , Vancouver, 11-13 August 2021, 3793-3810.
- Oest, A., Safaei, Y., Doupé, A., Ahn, G.J., et al . (2020) Sunrise to Sunset: Analyzing the End-to-End Life Cycle and Effectiveness of Phishing Attacks at Scale. 2020 29 th USENIX Security Symposium , Boston, 12-14 August 2020, 361-377.