Bài viết
Phát hiện lỗ hổng mã nguồn bằng cách sử dụng xử lý ngôn ngữ Nature và mạng đồ thị sâu
- Bùi Văn Công · Department of Information Technology, University of Economics and Technical Industries (VN)
- Đỗ Xuân Chợ (VN)
- Đỗ Trung Tuấn (VN)
Tóm tắt
Công nghiệp sản xuất phần mềm hưởng lợi từ các công cụ tự động sinh mã. Tuy nhiên cũng gặp thách thức về lỗ hổng phần mềm trong các mã sinh tự động đó. Liên quan đến phát hiện lỗ hổng phần mềm viết bằng các ngôn ngữ C và C++, bài báo đề xuất mô hình hỗn hợp giữa Graph Convolution Network (GCN) kết hợp với mô hình Bidirectional Encoder Representations from Transformers (BERT) và Dropout, gọi tắt là GBD. Trong pha 2 của mô hình GBD thực hiện đồng thời (i) trích xuất đặc trưng của đỉnh và cạnh dựa trên mạng tích chập đồ thị GCN; (ii) trích xuất đặc trưng đoạn mã sử dụng mô hình BERT; rồi thực hiện (iii) xây dựng hồ sơ mã nguồn, nhờ đồ thị đặc trưng mã CPG (Code Property Graph). Pha 3 của mô hình sẽ sử dụng kĩ thuật Dropout tránh quá khớp. Pha 4 của mô hình là bộ phân loại, xác định mã nguồn có lỗ hổng hay không. Các kết quả thực nghiệm cho thấy sự vượt trội của mô hình đề xuất so với các hướng tiếp cận khác trên tất cả các độ đo lần lượt đạt 61.21% và 88.94% tỷ lệ dự đoán đúng lỗ hổng mã nguồn và file bình thường. Đồng thời từ kết quả phân loại có thể thấy với token length là 512 thì mô hình GBD mang lại kết quả tốt, đồng đều nhất trên tất cả các độ đo Accuracy, Precision, Recall và F1 lần lượt đạt 86.65%, 38.59%, 66.21% và 48.76%. Phù hợp theo khảo sát của chúng tôi trên bộ dữ liệu thực nghiệm Verum vì có khoảng 70% các file mã nguồn có chiều dài nhỏ hơn 512 và lớn hơn 256. Ngoài ra mô hình GBD không chỉ hoạt động tốt trên một bộ dữ liệu mà còn có thể cho kết quả tốt trên nhiều bộ dữ liệu khác nhau. Cụ thể với bộ dữ liệu Verum, khi so sánh mô hình GBD với 5 hướng tiếp cận khác gồm REVEAL [1], Russell [2], VulDeePecker [3], SySeVR [4], Devign [5] cho kết quả tốt hơn lần lượt là 4% và từ 15% đến 57% trên 3 độ đo còn lại Precision, Recall, F1_score. Tương tự, khi so sánh GBD với hướng tiếp cận SySeVR [4] thì kết quả của GBD đã vượt trội hơn từ 3% đến 25% trên tất cả các độ đo. Còn với Devign [5] thì mô hình GBD mang lại hiệu quả hơn từ 5% đến 39% trên 3 độ đo Precision, Recall, F1_score. Với bộ dữ liệu FFmpeg+Qume, dựa trên kết quả thực nghiệm với độ đo Accuracy mô hình GBD hiệu quả hơn tất cả các nghiên cứu khác từ 0.2% đến 10%. Độ đo Precision, GBD cũng tốt hơn từ 0.3% đến 9% so với các hướng tiếp cận khác. Còn độ đo Recall thì GBD chỉ thấp hơn hướng tiếp cận REVEAL [1] khoảng 1.5% còn lại cao hơn tất cả các tiếp cận còn lại từ 10% đến hơn 31% còn độ đo Recall, độ đo F1-score của GBD cũng thấp hơn 0.3% so với REVEAL và cao hơn các nghiên cứu khác từ 7% đến 30%. Điều này chứng tỏ mô hình GBD không chỉ hiệu quả trên một bộ dữ liệu mà còn hiệu quả trên nhiều bộ dữ liệu khác nhau
Lượt tải theo tháng
Di chuột vào cột để xem số lượt tải.
Cách trích dẫn
Bùi Văn Công, Đỗ Xuân Chợ, Đỗ Trung Tuấn (2024). Phát hiện lỗ hổng mã nguồn bằng cách sử dụng xử lý ngôn ngữ Nature và mạng đồ thị sâu. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 3(23), 27-42. https://doi.org/10.54654/isj.v3i23.1057
Tài liệu tham khảo
- 1.S. Chakraborty, R. Krishna, Y. Ding and B. Ray, “Deep Learning based Vulnerability Detection: Are We There Yet?”, IEEE Transactions on Software Engineering, vol. 48, no. 9, pp. 3280-3296, 2022, doi: 10.1109/TSE.2021.3087402.
- 2.R. L. Russell, L. Kim, L. H. Hamilton, T. Lazovich, J. A. Harer, O. Ozdemir, P. M. Ellingwood and M. W. McConley, “Automated Vulnerability Detection in Source Code Using Deep Representation Learning”, In: 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), 2018; pp. 757-762, doi: 10.1109/ICMLA.2018.00120.
- 3.Z. Li, D. Zou, S. Xu, X. Ou, H. Jin, S. Wang, Z. Deng and Y. Zhong, “VulDeePecker: A Deep Learning-Based System for Vulnerability Detection”, in Network and Distributed Systems Security (NDSS) Symposium 2018, 18-21 February 2018, San Diego, CA, USA, https://arxiv. org/abs/1801.01681.
- 4.Z. Li, D. Zou, S. Xu, H. Jin, Y. Zhu and Z. Chen, "SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities", in IEEE Transactions on Dependable and Secure Computing, vol 19, no 4, pp 2244-2258, July-Aug. 2022, doi: 10.1109/TDSC.2021.3051525.
- 5.Y. Zhou, S. Liu, J. Siow, X. Du, and Y. Liu, “Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks”, in 33rd Int. Conf. Neural Inf. Process. Syst., Red Hook, NY, USA, no. 915, pp. 10197–10207, Dec. 2019.
- 6.NIST, Artificial intelligence (2022), , Access time VN 20/06/2024, https://www.nist.gov/.
- 7.B. Casey, J. C. S. Santos, G. Perry, “A Survey of Source Code Representations for Machine Learning-Based Cybersecurity Tasks”, ACM Comput. Surv. 37, vol. 4, no. 111, 35 pages, March. 2024. doi: 10.48500/arXiv.2403.10646.
- 8.CVE, All News, https://www.cve.org/Media/News/AllNews (2024), Access time: 20/06/2024.
- 9.CWE, CWE Top 25 Most Dangerous Software Weaknesses (2021), Access time: 20/06/2024, ,https://cwe.mitre.org/top25/archive/2021/2021_cwe_top25.html.
- 10.C. D. Xuan, D. H. Mai, M. C. Thanh and B. V. Cong, “A novel approach for software vulnerability detection based on intelligent cognitive computing”, the Journal of Supercomputing, vol 79, pp. 17042–17078, 2023. https://doi.org/10.1007/s11227-023-05282-4.
- 11.J. C. S. Santos, K. Tarrit and M. Mirakhorli, “A Catalog of Security Architecture Weaknesses”, Conference: 2017 IEEE International Conference on Software Architecture Workshops (ICSAW), 2017, pp. 220–223.
- 12.W. Cai, J. Chen, J. Yu and L. Gao, “A software vulnerability detection method based on deep learning with complex network analysis and subgraph partition”, in Information and Software Technology, vol. 164, no. 7, December. 2023, doi:https://doi.org/10.1016/j.infsof.2023.10732.
- 13.H. Wang, G. Ye, Z. Tang, S. H. Tan, S. Huang and D. Fang, “Combining Graph-Based Learning With Automated Data Collection for Code Vulnerability Detection”, in IEEE Transactions on Information Forensics and Security, vol. 16, pp. 1943-1958, 2021, doi: 10.1109/TIFS.2020.3044773.
- 14.H. Weic and M. Li, “Supervised deep features for software functional clone detection by exploiting lexical and syntactical information in source code”, in Proceedings of the TwentySixth International Joint Conference on Artificial Intelligence, Melbourne, Australia, pp. 3034–3040, August 2017.
- 15.X. Li, L. Wang, Y. Xin, Y. Yang, Q. Tang and Y. Chen, “Automated Software Vulnerability Detection Based on Hybrid Neural Network”, Appl. Sci. 2021, vol. 11, no. 7, pp. 3201. https://doi.org/10.3390/app11073201.
- 16.P. Zeng, G. Lin, L. Pan, Y. Tai and J. Zhang, “Software Vulnerability Analysis and Discovery Using Deep Learning Techniques: A Survey”, in IEEE Access, vol. 8, pp. 197158-197172, 2020, doi: 10.1109/ACCESS.2020.3034766.
- 17.V. K. Linh, N. V. Hung, T. N. Anh, D. D. Nhuan and D. C. Hien, “Enhance deep learning model for malware detection with a new image representation method”, the Journal of Science and Technology on Information security, vol. 21, no. 1, pp. 31-39, 2024, doi: https://doi.org/10.54654/isj.v1i21.1000.
- 18.F. Yamaguchi, N. Golde, D. Arp and K. Rieck, “Modeling and Discovering Vulnerabilities with Code Property Graphs”, IEEE Symposium on Security and Privacy, Berkeley, CA, USA, 2014, doi: 10.1109/SP.2014.44
- 19.J. Devlin,M. W. Chang, K. Lee and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2018, pp. 4171-4186, arXiv:1810.04805.
- 20.N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique”, Journal of artificial intelligence research, vol. 16, pp. 321–357, 2002.
- 21.B. Liu, W. Guan, C. Yang, Z. Fang and Z. Lu, “Transformer and Graph Convolutional Network for Text Classification”, International Journal of Computational Intelligence Systems, vol. 16, October 2023. https://doi.org/10.1007/s44196-023-00337-z.
- 22.F. Subhan, X. Wu, L. Bo, X. Sun and M. Rahman, “A deep learning‐based approach for software vulnerability detection using code metrics”, Institution of Engineering and Technology - IET, vol. 16, pp. 516-526, 2022.
- 23.V. C. Bui and X. C. Do, “Detecting software vulnerabilities based on source code analysis using GCN transformer”, in 2023 RIVF Int. Conf. Comput. Commun. Technol. (RIVF), pp.112–117, 2023.
- 24.G. Tang, L. Yang, L. Zhang, W. Cao, L. Meng, H. He, H. Kuang, F. Yang and H. Wang, “An attention-based automatic vulnerability detection approach with GGNN”, Int. J. Mach. Learn. & Cyber, vol. 14, pp. 3113–3127, 2023, https://doi.org/10.1007/s13042-023-01824-7
- 25.Download Ffmpeg, Access time: 20/06/2024, https://ffmpeg.org/download.html.
- 26.T. T. Nguyen and H. D. Vo, “Context-based statement-level vulnerability localization”, Information and Software Technology, vol. 169, 107406 pages, 2024.
- 27.JOERN, The Bug Hunter's Workbench (2024), , Access time: 20/06/2024, https://joern.io/.
- 28.B. Chernis and R. Verma, “Machine Learning Methods for Software Vulnerability Detection”, in IWSPA '18: Proceedings of the Fourth ACM International Workshop on Security and Privacy Analytics, March 19–21, 2018, Tempe, AZ, USA, pp. 31-39. https://doi.org/10.1145/3180445.3180453.
- 29.Q. Li, J. Song, D. Tan, H. Wang and J. Liu, “PDGraph: A Large-Scale Empirical Study on Project Dependency of Security Vulnerabilities”, in 2021 51st Annual IEEE/IFIP Int. Conf. Depen. Sys. Net. (DSN), pp.161–173, 2021.
- 30.T. N. Kipf and M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks”, International Conference on Learning Representations, 9 September 2016, doi: 10.48550/arXiv.1609.02907.
- 31.K. Yang, P. Miller and J. Martinez-Del-Rincon, “Convolutional Neural Network for Software Vulnerability Detection”, in IEEE Transactions on Information Forensics and Security, 2022, DOI: 10.1109/Cyber-CI55324.2022.10032684
- 32.J. Chen, Y. Yin, S. Cai, W. Wang, S. Wang and J. Chen, “iGnnVD: A novel software vulnerability detection model based on integrated graph neural networks”, Science of Computer Programming, vol. 238, pp. 103156, 2024.
- 33.H. Wang, Z. Qu and L. Sun, “E-GVD: Efficient Software Vulnerability Detection Techniques Based on Graph Neural Network”, ICST Transactions on Scalable Information Systems, vol. 11, March 2024. doi:10.4108/eetsis.5056
- 34.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever and R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research, vol. 15, no. 56, pp. 1929-1958, 2014.
- 35.P. Baldi and P. J. Sadowski, “Understanding Dropout”, In: Proceedings in the Advances in Neural Information Processing Systems 26. Red Hook, NY, USA, December. 2013.
- 36.X. Li, S. Chen, X. Hu and J. Yang, “Understanding the Disharmony Between Dropout and Batch Normalization by Variance Shift”. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019; pp. 2677-2685, doi. 10.1109/CVPR.2019.00279.
- 37.K. Duan, S. S. Keerthi, W. Chu, S. K. Shevade and A. N. Poo, “Multi-category Classification by Soft-Max Combination of Binary Classifiers”, In proceedings of the 4th International Workshop, MCS 2003 Guildford, UK, 11–13, pp 125–134, June 2003. doi: 10.1007/3-540-44938-8_13.
- 38.X. Xu, C. Liu, Q. Feng, H. Yin, L. Song and D. Song, “Neural networkbased graph embedding for cross-platform binary code similarity detection”, in Proc. ACM SIGSAC Conf. Comput. Commun. Secur, pp. 363–376, Oct. 2017.
- 39.Y. Li, S. Wang and T. N. Nguyen, “Vulnerability detection with fine-grained interpretations”, Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp. 292–303, August 2021. https://doi.org/10.1145/3468264.3468597.