Bài viết
Generating evasive payloads for assessing Web Application Firewalls with Reinforcement Learning and Pre-trained Language Models
- Trần Gia Bảo · University of Information Technology, VNU-HCM (VN)
- Đinh Công Đức · University of Information Technology, VNU-HCM (VN)
- Phan Thế Duy · University of Information Technology, VNU-HCM (VN)
Tóm tắt
Tường lửa ứng dụng Web (WAF) đóng vai trò là một cơ chế phòng thủ quan trọng chống lại nhiều dạng tấn công dựa trên web như SQL Injection (SQLi), Cross-Site Scripting (XSS), Server-Side Request Forgery (SSRF), Remote Code Execution (RCE) và NoSQL Injection. Tuy nhiên, các tin tặc hiện đại, tinh vi thường tạo ra các payload được làm rối và né tránh nhằm vượt qua các luật WAF truyền thống. Để đánh giá và thách thức hiệu quả độ bền vững của WAF, Để đánh giá và thách thức hiệu quả độ mạnh mẽ của các Tường lửa ứng dụng web (WAF), nhóm tắc giả đề xuất DEG-WAF, một khung công tác Tạo sinh Payload né tránh sâu (Deep Evasion Generation) tận dụng mô hình ngôn ngữ Lớn (LLM) kết hợp với học tăng cường (RL) để tạo ra các payload né tránh nhằm chống lại các WAF. Hệ thống bao gồm bốn thành phần chính: một tác tử sinh payload dựa trên LLM đã được huấn luyện trước (OPT-125M), một mô hình thưởng mô phỏng hành vi của WAF, một tác tử lấy mẫu dựa trên ngữ pháp để đảm bảo tính hợp lệ cú pháp, và một tác tử RL được huấn luyện với thuật toán Proximal Policy Optimization (PPO) hoặc Advantage ActorCritic (A2C) nhằm tinh chỉnh chiến lược sinh. Các đánh giá thực nghiệm trên các WAF thực tế, bao gồm ModSecurity và SafeLine, cho thấy mô hình dựa trên A2C vượt trội đáng kể so với các LLM nguyên bản — đạt tỷ lệ né tránh 80.16% với SQLi và 74.70% với NoSQLi trên ModSecurity, và 97.8% với RCE trên SafeLine. Những kết quả này nhấn mạnh tiềm năng của khung LLM-RL do chúng tôi đề xuất như một nền tảng vững chắc để đánh giá và nâng cao khả năng chống chịu của các hệ thống WAF trong điều kiện đối kháng.
Lượt tải theo tháng
Di chuột vào cột để xem số lượt tải.
Cách trích dẫn
Trần Gia Bảo, Đinh Công Đức, Phan Thế Duy (2025). Generating evasive payloads for assessing Web Application Firewalls with Reinforcement Learning and Pre-trained Language Models. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 2(25), 78-96. https://doi.org/10.54654/isj.v2i25.1128
Tiểu sử tác giả
Trần Gia Bảo, University of Information Technology, VNU-HCM
Quá trình đào tạo: Sinh viên năm cuối chuyên ngành An toàn Thông tin (Chương trình Cử nhân Tài năng)Hướng nghiên cứu hiện nay: An toàn ứng dụng web, kiểm thử xâm nhập, học tăng cường trong bảo mật, mô hình ngôn ngữ lớn, phát hiện lỗi bảo mật trong mã nguồn, và diễn giải mô hình AI (XAI).Đinh Công Đức, University of Information Technology, VNU-HCM
Quá trình đào tạo: Sinh viên năm cuối chuyên ngành An toàn Thông tin (Chương trình Cử nhân Tài năng)Hướng nghiên cứu hiện nay: An toàn ứng dụng web, kiểm thử xâm nhập, học tăng cường trong bảo mật, mô hình ngôn ngữ lớn.Phan Thế Duy, University of Information Technology, VNU-HCM
Quá trình đào tạo: Tiến sĩ ngành Công nghệ thông tin, chuyên ngành An toàn thông tinHướng nghiên cứu hiện nay: Kiểm thử thâm nhập, an toàn ứng dụng web, bảo mật Hợp đồng thông minh, phân tích mã độc.
Tài liệu tham khảo
- 1.O. Fredj, O. Cheikhrouhou, M. Krichen, H. Hamam, and A. Derhab, “An owasp top ten driven survey on web application protection methods,” 11 2020.
- 2.V. Clincy and H. Shahriar, “Web application firewall: Network security models and configuration,” in 2018 IEEE 42nd Annual Computer Software and Applications Conference (COMPSAC), vol. 01, 2018, pp. 835–836. DOI: 10.1109/COMPSAC.2018.00144
- 3.A. Coscia, V. Dentamaro, S. Galantucci, A. Maci, and G. Pirlo, “Progesi: a proxy grammar to enhance web application firewall for sql injection prevention,” IEEE Access, vol. 12, pp. 107 689–107 703, 08 2024. DOI:
- 4.1109/ACCESS.2024.3438092
- 5.N. N. Thanh, V.-G. Ung, P. T. Duy, and V.-H. Pham, “A study on adversarial attacks for benchmarking deep learning-based
- 6.web application firewalls,” in 2024 RIVF International Conference on Computing and Communication Technologies (RIVF). IEEE, 2024, pp. 151–155.
- 7.A. Valenza, L. Demetrio, G. Costa, and G. Lagorio, “Waf-a-mole: An adversarial tool for assessing ml-based wafs,” SoftwareX, vol. 11, p. 100367, 2020. DOI: https://doi.org/10.1016/j.softx.2019.100367
- 8.H. Liang, X. Li, D. Xiao, J. Liu, Y. Zhou, A. Wang, and J. Li, “Generative pre-trained transformer-based reinforcement learning for testing web application firewalls,” IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 1, pp. 309–324, 2024. DOI: 10.1109/TDSC.2023.3252523
- 9.D. Leung, O. Tsai, K. Hashemi, B. Tayebi, and M. A. Tayebi, “Xploitsql: Advancing adversarial sql injection attack generation with language models and reinforcement learning,” in Proceedings of the 33rd ACM
- 10.International Conference on Information and Knowledge Management, ser. CIKM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 4653–4660. DOI: 10.1145/3627673.3680102
- 11.S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao, “Large language models: A survey,” . , https://arxiv.org/abs/2402.06196
- 12.H. Xu, S. Wang, N. Li, K. Wang, Y. Zhao, K. Chen, T. Yu, Y. Liu, and H. Wang, “Large language models for cyber security: A
- 13.systematic literature review,” arXiv preprint arXiv:2405.04760, 2024.
- 14.V. Babaey and A. Ravindran, “Gensqli: A generative artificial intelligence framework for automatically securing web application firewalls against structured query language injection attacks,” Future Internet, vol. 17, no. 1, 2025. DOI: 10.3390/fi17010008
- 15.D. Miczek, D. Gabbireddy, and S. Saha, “Leveraging llm to strengthen ml-based cross-site scripting detection,” . , https: //arxiv.org/abs/2504.21045
- 16.Z. Gui, E. Wang, B. Deng, M. Zhang, Y. Chen, S. Wei, W. Xie, and B. Wang, “Sqligpt: Evaluating and utilizing large language models for automated sql injection black-box detection,” Applied Sciences, vol. 14, no. 16, 2024. DOI: 10.3390/app14166929
- 17.V. Babaey and A. Ravindran, “Genxss: an ai-driven framework for automated detection of xss attacks in wafs,” . , https://arxiv.org/ abs/2504.08176
- 18.H. Kheddar, D. W. Dawoud, A. I. Awad, Y. Himeur, and M. K. Khan, “Reinforcementlearning-based intrusion detection in communication networks: A review,” IEEE Communications Surveys & Tutorials, 2024.
- 19.M. Ghasemi and D. Ebrahimi, “Introduction to reinforcement learning,” . , https://arxiv.
- 20.org/abs/2408.07712
- 21.S. Finistrella, S. Mariani, and F. Zambonelli, “Multi-agent reinforcement learning for cybersecurity: Classification and survey,” Intelligent Systems with Applications, p. 200495, 2025.
- 22.C. Folini and I. Ristic, ModSecurity Handbook, Second Edition, 2nd ed. London, GBR: Feisty Duck, 2017.
- 23.Chaitin Technology, “Safeline web application firewall documentation,” 2023. , Access date: 10/6/2025, https://docs.waf .chaitin.com
- 24.Wargio, “Naxsi: Nginx anti xss and sql injection,” 2023. , Access date: 10/6/2025, https://github.com/wargio/naxsi
- 25.OWASP ModSecurity Core Rule Set Team, “Owasp core rule set (crs),” 2023. , Access date: 9/6/2025, https://github.com/corerules et/coreruleset
- 26.S. Dhote, A. Magdum, S. Singh, and D. Raigar, “Ml based web application firewall for signature and anomaly detection using feature extraction,” in 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT), 2024, pp. 1–6. DOI: 10.1109/ICCCNT61001.2024.10725511
- 27.K. S. Kalyan, “A survey of gpt-3 family large language models including chatgpt and gpt-4,” Natural Language Processing Journal, vol. 6, p. 100048, 2024. DOI: https://doi.org/10.1016/j.nlp.2023.100048
- 28.G. Yenduri, R. M, C. S. G, S. Y, G. Srivastava, P. K. R. Maddikunta, D. R. G, R. H. Jhaveri, P. B, W. Wang, A. V. Vasilakos,
- 29.and T. R. Gadekallu, “Generative pre-trained transformer: A comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions” , https://arxiv.org/abs/2305.10435
- 30.S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer, “Opt: Open pre-trained transformer language models”, https://arxiv.org/abs/2205.01068
- 31.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar,
- 32.A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models”, https://arxiv.org/abs/2302.13971
- 33.H. Zhou, C. Hu, Y. Yuan, Y. Cui, Y. Jin, C. Chen, H. Wu, D. Yuan, L. Jiang, D. Wu, X. Liu, J. Zhang, X. Wang, and J. Liu, “Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,” IEEE Communications Surveys Tutorials, vol. 27, no. 3, pp. 1955–2005, 2025. DOI: 10.1109/COMST.2024.3465447
- 34.Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,” High-Confidence Computing, vol. 4, no. 2, p. 100211, Jun. 2024. DOI: 10.1016/j.hcc.2024.100211
- 35.J. Zhang, H. Bu, H. Wen, Y. Liu, H. Fei, R. Xi, L. Li, Y. Yang, H. Zhu, and D. Meng, “When llms meet cybersecurity: A systematic literature review,” Cybersecurity, vol. 8, no. 1, p. 55, 2025.
- 36.A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret, “Warm: On the benefits
- 37.of weight averaged reward models,” . , https://arxiv.org/abs/2401.12187
- 38.V. Atlidakis, R. Geambasu, P. Godefroid, M. Polishchuk, and B. Ray, “Pythia: Grammar-based fuzzing of rest apis with coverage-guided feedback and learning-based mutations,” arXiv preprint arXiv:2005.11498, 2020.
- 39.P. Godefroid, H. Peleg, and R. Singh, “Learn&fuzz: Machine learning for
- 40.input fuzzing,” in 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2017, pp. 50–59.
- 41.P. Srivastava and M. Payer, “Gramatron: effective grammar-aware fuzzing,” in Proceedings of the 30th ACM SIGSOFT
- 42.International Symposium on Software Testing and Analysis, ser. ISSTA 2021. New York, NY, USA: Association for Computing Machinery, 2021, p. 244–256. DOI: 10.1145/3460319.3464814
- 43.Z. Qu, X. Ling, T. Wang, X. Chen, S. Ji, and C. Wu, “Advsqli: Generating adversarial sql injections against realworld waf-as-a-service,” IEEE Transactions on Information Forensics and Security, vol. 19, p. 2623–2638, 2024. DOI: 10.1109/tifs.2024.3350911
- 44.K. Li, H. Yang, and W. Visser, “Evolutionary multi-task injection testing on web application firewalls,” . , https://arxiv.org/abs/2206.05743
- 45.C. Wu, J. Chen, S. Zhu, W. Feng, R. Du, and Y. Xiang, “Wafbooster: Automatic boosting of waf security against mutated malicious payloads,” . , https://arxiv.org/abs/2501.140 08
- 46.F. Yang, W. Zhou, Z. Liu, D. Zhao, and D. Held, “Reinforcement learning in a safetyembedded mdp with trajectory optimizatioz , https://arxiv.org/abs/2310.06903.
- 47.M.-A. Chadi and H. Mousannif, “Understanding reinforcement learning algorithms: The progress from basic qlearning to proximal policy optimization”, https://arxiv.org/abs/2304.00026
- 48.OpenAI, “Spinning up - proximal policy optimization (ppo),” 2024. , Access date: 10/6/2025, https://spinningup.openai.com/en /latest/algorithms/ppo.html
- 49.GeeksforGeeks Contributors, “Actor-critic algorithm in reinforcement learning,” 2023. , Access date: 10/6/2025, https://www.geeksf orgeeks.org/machine-learning/actor-critic-a lgorithm-in-reinf orcement-learning/
- 50.Swisskyrepo, “Payloads all the things,” 2023, Access date: 10/6/2025, https://github.com/swisskyrepo/PayloadsAllTheThings.