Bài viết

Generating evasive payloads for assessing Web Application Firewalls with Reinforcement Learning and Pre-trained Language Models

Từ khóaTường lửa ứng dụng webhọc tăng cườngngôn ngữ lớntác tử sinh payload

Tóm tắt

 Tường lửa ứng dụng Web (WAF) đóng vai trò là một cơ chế phòng thủ quan trọng chống lại nhiều dạng tấn công dựa trên web như SQL Injection (SQLi), Cross-Site Scripting (XSS), Server-Side Request Forgery (SSRF), Remote Code Execution (RCE) và NoSQL Injection. Tuy nhiên, các tin tặc hiện đại, tinh vi thường tạo ra các payload được làm rối và né tránh nhằm vượt qua các luật WAF truyền thống. Để đánh giá và thách thức hiệu quả độ bền vững của WAF, Để đánh giá và thách thức hiệu quả độ mạnh mẽ của các Tường lửa ứng dụng web (WAF), nhóm tắc giả đề xuất DEG-WAF, một khung công tác Tạo sinh Payload né tránh sâu (Deep Evasion Generation) tận dụng mô hình ngôn ngữ Lớn (LLM) kết hợp với học tăng cường (RL) để tạo ra các payload né tránh nhằm chống lại các WAF. Hệ thống bao gồm bốn thành phần chính: một tác tử sinh payload dựa trên LLM đã được huấn luyện trước (OPT-125M), một mô hình thưởng mô phỏng hành vi của WAF, một tác tử lấy mẫu dựa trên ngữ pháp để đảm bảo tính hợp lệ cú pháp, và một tác tử RL được huấn luyện với thuật toán Proximal Policy Optimization (PPO) hoặc Advantage ActorCritic (A2C) nhằm tinh chỉnh chiến lược sinh. Các đánh giá thực nghiệm trên các WAF thực tế, bao gồm ModSecurity và SafeLine, cho thấy mô hình dựa trên A2C vượt trội đáng kể so với các LLM nguyên bản — đạt tỷ lệ né tránh 80.16% với SQLi và 74.70% với NoSQLi trên ModSecurity, và 97.8% với RCE trên SafeLine. Những kết quả này nhấn mạnh tiềm năng của khung LLM-RL do chúng tôi đề xuất như một nền tảng vững chắc để đánh giá và nâng cao khả năng chống chịu của các hệ thống WAF trong điều kiện đối kháng.

Lượt tải theo tháng

0418110/2511/2512/2501/2602/2603/2604/2605/26

Di chuột vào cột để xem số lượt tải.

Cách trích dẫn

Trần Gia Bảo, Đinh Công Đức, Phan Thế Duy (2025). Generating evasive payloads for assessing Web Application Firewalls with Reinforcement Learning and Pre-trained Language Models. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 2(25), 78-96. https://doi.org/10.54654/isj.v2i25.1128

Tiểu sử tác giả

  • Trần Gia Bảo, University of Information Technology, VNU-HCM

    Quá trình đào tạo: Sinh viên năm cuối chuyên ngành An toàn Thông tin (Chương trình Cử nhân Tài năng) 
    Hướng nghiên cứu hiện nay: An toàn ứng dụng web, kiểm thử xâm nhập, học tăng cường trong bảo mật, mô hình ngôn ngữ lớn, phát hiện lỗi bảo mật trong mã nguồn, và diễn giải mô hình AI (XAI).
  • Đinh Công Đức, University of Information Technology, VNU-HCM

    Quá trình đào tạo: Sinh viên năm cuối chuyên ngành An toàn Thông tin (Chương trình Cử nhân Tài năng) 
    Hướng nghiên cứu hiện nay: An toàn ứng dụng web, kiểm thử xâm nhập, học tăng cường trong bảo mật, mô hình ngôn ngữ lớn. 
  • Phan Thế Duy, University of Information Technology, VNU-HCM

    Quá trình đào tạo: Tiến sĩ ngành Công nghệ thông tin, chuyên ngành An toàn thông tin
    Hướng nghiên cứu hiện nay: Kiểm thử thâm nhập, an toàn ứng dụng web, bảo mật Hợp đồng thông minh, phân tích mã độc.

Tài liệu tham khảo

  1. 1.O. Fredj, O. Cheikhrouhou, M. Krichen, H. Hamam, and A. Derhab, “An owasp top ten driven survey on web application protection methods,” 11 2020.
  2. 2.V. Clincy and H. Shahriar, “Web application firewall: Network security models and configuration,” in 2018 IEEE 42nd Annual Computer Software and Applications Conference (COMPSAC), vol. 01, 2018, pp. 835–836. DOI: 10.1109/COMPSAC.2018.00144
  3. 3.A. Coscia, V. Dentamaro, S. Galantucci, A. Maci, and G. Pirlo, “Progesi: a proxy grammar to enhance web application firewall for sql injection prevention,” IEEE Access, vol. 12, pp. 107 689–107 703, 08 2024. DOI:
  4. 4.1109/ACCESS.2024.3438092
  5. 5.N. N. Thanh, V.-G. Ung, P. T. Duy, and V.-H. Pham, “A study on adversarial attacks for benchmarking deep learning-based
  6. 6.web application firewalls,” in 2024 RIVF International Conference on Computing and Communication Technologies (RIVF). IEEE, 2024, pp. 151–155.
  7. 7.A. Valenza, L. Demetrio, G. Costa, and G. Lagorio, “Waf-a-mole: An adversarial tool for assessing ml-based wafs,” SoftwareX, vol. 11, p. 100367, 2020. DOI: https://doi.org/10.1016/j.softx.2019.100367
  8. 8.H. Liang, X. Li, D. Xiao, J. Liu, Y. Zhou, A. Wang, and J. Li, “Generative pre-trained transformer-based reinforcement learning for testing web application firewalls,” IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 1, pp. 309–324, 2024. DOI: 10.1109/TDSC.2023.3252523
  9. 9.D. Leung, O. Tsai, K. Hashemi, B. Tayebi, and M. A. Tayebi, “Xploitsql: Advancing adversarial sql injection attack generation with language models and reinforcement learning,” in Proceedings of the 33rd ACM
  10. 10.International Conference on Information and Knowledge Management, ser. CIKM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 4653–4660. DOI: 10.1145/3627673.3680102
  11. 11.S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao, “Large language models: A survey,” . , https://arxiv.org/abs/2402.06196
  12. 12.H. Xu, S. Wang, N. Li, K. Wang, Y. Zhao, K. Chen, T. Yu, Y. Liu, and H. Wang, “Large language models for cyber security: A
  13. 13.systematic literature review,” arXiv preprint arXiv:2405.04760, 2024.
  14. 14.V. Babaey and A. Ravindran, “Gensqli: A generative artificial intelligence framework for automatically securing web application firewalls against structured query language injection attacks,” Future Internet, vol. 17, no. 1, 2025. DOI: 10.3390/fi17010008
  15. 15.D. Miczek, D. Gabbireddy, and S. Saha, “Leveraging llm to strengthen ml-based cross-site scripting detection,” . , https: //arxiv.org/abs/2504.21045
  16. 16.Z. Gui, E. Wang, B. Deng, M. Zhang, Y. Chen, S. Wei, W. Xie, and B. Wang, “Sqligpt: Evaluating and utilizing large language models for automated sql injection black-box detection,” Applied Sciences, vol. 14, no. 16, 2024. DOI: 10.3390/app14166929
  17. 17.V. Babaey and A. Ravindran, “Genxss: an ai-driven framework for automated detection of xss attacks in wafs,” . , https://arxiv.org/ abs/2504.08176
  18. 18.H. Kheddar, D. W. Dawoud, A. I. Awad, Y. Himeur, and M. K. Khan, “Reinforcementlearning-based intrusion detection in communication networks: A review,” IEEE Communications Surveys & Tutorials, 2024.
  19. 19.M. Ghasemi and D. Ebrahimi, “Introduction to reinforcement learning,” . , https://arxiv.
  20. 20.org/abs/2408.07712
  21. 21.S. Finistrella, S. Mariani, and F. Zambonelli, “Multi-agent reinforcement learning for cybersecurity: Classification and survey,” Intelligent Systems with Applications, p. 200495, 2025.
  22. 22.C. Folini and I. Ristic, ModSecurity Handbook, Second Edition, 2nd ed. London, GBR: Feisty Duck, 2017.
  23. 23.Chaitin Technology, “Safeline web application firewall documentation,” 2023. , Access date: 10/6/2025, https://docs.waf .chaitin.com
  24. 24.Wargio, “Naxsi: Nginx anti xss and sql injection,” 2023. , Access date: 10/6/2025, https://github.com/wargio/naxsi
  25. 25.OWASP ModSecurity Core Rule Set Team, “Owasp core rule set (crs),” 2023. , Access date: 9/6/2025, https://github.com/corerules et/coreruleset
  26. 26.S. Dhote, A. Magdum, S. Singh, and D. Raigar, “Ml based web application firewall for signature and anomaly detection using feature extraction,” in 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT), 2024, pp. 1–6. DOI: 10.1109/ICCCNT61001.2024.10725511
  27. 27.K. S. Kalyan, “A survey of gpt-3 family large language models including chatgpt and gpt-4,” Natural Language Processing Journal, vol. 6, p. 100048, 2024. DOI: https://doi.org/10.1016/j.nlp.2023.100048
  28. 28.G. Yenduri, R. M, C. S. G, S. Y, G. Srivastava, P. K. R. Maddikunta, D. R. G, R. H. Jhaveri, P. B, W. Wang, A. V. Vasilakos,
  29. 29.and T. R. Gadekallu, “Generative pre-trained transformer: A comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions” , https://arxiv.org/abs/2305.10435
  30. 30.S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer, “Opt: Open pre-trained transformer language models”, https://arxiv.org/abs/2205.01068
  31. 31.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar,
  32. 32.A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models”, https://arxiv.org/abs/2302.13971
  33. 33.H. Zhou, C. Hu, Y. Yuan, Y. Cui, Y. Jin, C. Chen, H. Wu, D. Yuan, L. Jiang, D. Wu, X. Liu, J. Zhang, X. Wang, and J. Liu, “Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,” IEEE Communications Surveys Tutorials, vol. 27, no. 3, pp. 1955–2005, 2025. DOI: 10.1109/COMST.2024.3465447
  34. 34.Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,” High-Confidence Computing, vol. 4, no. 2, p. 100211, Jun. 2024. DOI: 10.1016/j.hcc.2024.100211
  35. 35.J. Zhang, H. Bu, H. Wen, Y. Liu, H. Fei, R. Xi, L. Li, Y. Yang, H. Zhu, and D. Meng, “When llms meet cybersecurity: A systematic literature review,” Cybersecurity, vol. 8, no. 1, p. 55, 2025.
  36. 36.A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret, “Warm: On the benefits
  37. 37.of weight averaged reward models,” . , https://arxiv.org/abs/2401.12187
  38. 38.V. Atlidakis, R. Geambasu, P. Godefroid, M. Polishchuk, and B. Ray, “Pythia: Grammar-based fuzzing of rest apis with coverage-guided feedback and learning-based mutations,” arXiv preprint arXiv:2005.11498, 2020.
  39. 39.P. Godefroid, H. Peleg, and R. Singh, “Learn&fuzz: Machine learning for
  40. 40.input fuzzing,” in 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2017, pp. 50–59.
  41. 41.P. Srivastava and M. Payer, “Gramatron: effective grammar-aware fuzzing,” in Proceedings of the 30th ACM SIGSOFT
  42. 42.International Symposium on Software Testing and Analysis, ser. ISSTA 2021. New York, NY, USA: Association for Computing Machinery, 2021, p. 244–256. DOI: 10.1145/3460319.3464814
  43. 43.Z. Qu, X. Ling, T. Wang, X. Chen, S. Ji, and C. Wu, “Advsqli: Generating adversarial sql injections against realworld waf-as-a-service,” IEEE Transactions on Information Forensics and Security, vol. 19, p. 2623–2638, 2024. DOI: 10.1109/tifs.2024.3350911
  44. 44.K. Li, H. Yang, and W. Visser, “Evolutionary multi-task injection testing on web application firewalls,” . , https://arxiv.org/abs/2206.05743
  45. 45.C. Wu, J. Chen, S. Zhu, W. Feng, R. Du, and Y. Xiang, “Wafbooster: Automatic boosting of waf security against mutated malicious payloads,” . , https://arxiv.org/abs/2501.140 08
  46. 46.F. Yang, W. Zhou, Z. Liu, D. Zhao, and D. Held, “Reinforcement learning in a safetyembedded mdp with trajectory optimizatioz , https://arxiv.org/abs/2310.06903.
  47. 47.M.-A. Chadi and H. Mousannif, “Understanding reinforcement learning algorithms: The progress from basic qlearning to proximal policy optimization”, https://arxiv.org/abs/2304.00026
  48. 48.OpenAI, “Spinning up - proximal policy optimization (ppo),” 2024. , Access date: 10/6/2025, https://spinningup.openai.com/en /latest/algorithms/ppo.html
  49. 49.GeeksforGeeks Contributors, “Actor-critic algorithm in reinforcement learning,” 2023. , Access date: 10/6/2025, https://www.geeksf orgeeks.org/machine-learning/actor-critic-a lgorithm-in-reinf orcement-learning/
  50. 50.Swisskyrepo, “Payloads all the things,” 2023, Access date: 10/6/2025, https://github.com/swisskyrepo/PayloadsAllTheThings.

Bài viết liên quan