Bài viết

Đề xuất phương pháp tổng hợp mới để tìm kiếm tin tức về hack theo ngữ nghĩa

Từ khóatìm kiếm ngữ nghĩamô hình ngôn ngữ lớntin tức về hoạt động tấn công

Tóm tắt

Tìm kiếm thông tin về các hoạt động tấn công một cách hiệu quả là chủ đề được thảo luận sôi nổi trong những năm gần đây. Nhiều thách thức gặp phải khi tìm kiếm những thông tin này. Đặc biệt, những khó khăn có thể gặp phải khi hiểu các thuật ngữ, ý tưởng, công cụ và một số mục không phổ biến chỉ dành riêng cho việc tấn công. Việc hiểu hiệu quả các từ đồng nghĩa và đa nghĩa là cần thiết. Những lý do này đóng vai trò là động lực thúc đẩy nỗ lực của nhóm tác giả nhằm phát triển một phương pháp hiệu quả để tìm kiếm thông tin hack ngữ nghĩa. Việc áp dụng các kỹ thuật xử lý ngôn ngữ tự nhiên hiện đại vào tìm kiếm theo ngữ nghĩa đã cải thiện đáng kể việc truy xuất thông tin bằng cách nâng cao độ chính xác và tính liên quan của kết quả tìm kiếm. So với các phương pháp truyền thống, các mô hình mạng nơron xử lý hiệu quả các từ đồng nghĩa và đa nghĩa. Tuy nhiên, mô hình càng lớn thì thời gian xử lý càng nhiều. Bài báo này đề xuất phương pháp tìm kiếm ngữ nghĩa (NESS) bằng cách kết hợp các mô hình nhúng nhỏ. Khi đánh giá trên tập dữ liệu chứa hơn 300.000 bản ghi từ trang Hacker News, NESS cải thiện đáng kể chất lượng xếp hạng và độ chính xác truy xuất so với các kỹ thuật hiện có, đồng thời thời gian xử lý giảm một nửa so với mô hình lớn có độ chính xác tốt nhất. Kết quả nhấn mạnh việc cân nhắc giữa độ phức tạp của mô hình, độ chính xác của kết quả và hiệu quả truy vấn, cung cấp những hiểu biết để tối ưu hóa các hệ thống tìm kiếm theo ngữ nghĩa.

Lượt tải theo tháng

0101910/2411/2412/2403/2504/2505/2506/2507/2508/2509/2510/2511/2512/2501/2602/2603/2604/2605/26

Di chuột vào cột để xem số lượt tải.

Cách trích dẫn

Đỗ Ngọc Long, Nguyễn Thế Hùng, Nguyễn Trung Dũng, Đỗ Văn Khánh, Nguyễn Anh Tú, Phạm Thị Bích Vân (2024). Đề xuất phương pháp tổng hợp mới để tìm kiếm tin tức về hack theo ngữ nghĩa. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 2(22), 83-92. https://doi.org/10.54654/isj.v2i22.1033

Tài liệu tham khảo

  1. 1.Sun, Nan, et al (2023). Cyber threat intelligence mining for proactive cybersecurity defense: a survey and new perspectives. IEEE Communications Surveys & Tutorials 2023.
  2. 2.Thakur, Manikant. Cyber security threats and countermeasures in digital age. Journal of Applied Science and Education (JASE) 4.1 (2024): 1-20.
  3. 3.Benjamin, Victor, and Hsinchun Chen. "Developing understanding of hacker language through the use of lexical semantics." 2015 IEEE International Conference on Intelligence and Security Informatics (ISI). IEEE, 2015.
  4. 4.Li, Ying, et al. "NEDetector: Automatically extracting cybersecurity neologisms from hacker forums." Journal of Information Security and Applications 58 (2021): 102784.
  5. 5.Satyapanich, Taneeya, Tim Finin, and Francis Ferraro. "Extracting rich semantic information about cybersecurity events." 2019 IEEE International Conference on Big Data (Big Data). IEEE, 2019.
  6. 6.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. CoRR, abs/1301.3781.
  7. 7.Mitra, B., & Craswell, N. (2018). An introduction to neural information retrieval. Foundations and Trends in Information Retrieval, 13(1), 1-126.
  8. 8.Mehrish, A., Majumder, N., Bharadwaj, R., Mihalcea, R., & Poria, S. (2023). A review of deep learning techniques for speech processing. Information Fusion, 101869.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805.
  10. 10.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. Technical report, OpenAI.
  11. 11.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (pp. 5998-6008).
  12. 12.Hugo Touvron, Thibaut Lavril, Xavier Martinet, et al. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
  13. 13.Tom Kenter, Alexey Borisov, and Maarten de Rijke. 2016. Siamese CBOW: Optimizing word embeddings for sentence representations. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 941-951).
  14. 14.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7, 249-266.
  15. 15.Brandon Nye, JinJin Li, Ramya Patel, et al. 2018. A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) (pp. 197-203).
  16. 16.Jiafeng Guo, Yixing Fan, Qingyao Ai, and W. Bruce Croft. 2016. A deep relevance matching model for ad-hoc retrieval. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 2251-2259).
  17. 17.JD Prater. AI-Powered Semantic Search: Everything You Need to Know. https://www.graft.com/blog/the-future-is-semantic-transforming-search-in-the-age-of-ai (2023).
  18. 18.Trotman, Andrew, Antti Puurula, and Blake Burgess (2014). Improvements to BM25 and language models examined. Proceedings of the 2014 Australasian Document Computing Symposium (ADCS 2014) (pp. 58-65).
  19. 19.Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval 3, 4 (2009), 333-389.
  20. 20.Zhu, Y., Yuan, H., Wang, S., Liu, J., Liu, W., Deng, C., Dou, Z., & Wen, J. 2023. Large Language Models for Information Retrieval: A Survey. arXiv, abs/2308.07107.
  21. 21.Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021. MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2140–2151, Online. Association for Computational Linguistics.
  22. 22.Reimers, Nils and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 3982-3992).
  23. 23.Khattab, Omar, and Matei Zaharia. Colbert: Efficient and effective passage search via contextualized late interaction over bert. Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 2020 (pp. 39–48).
  24. 24.Sanh, Victor et al. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv abs/1910.01108.
  25. 25.Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., & Liu, Z. 2024. BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv, abs/2402.03216.
  26. 26.Ryan Michael. kerinin/hackernews-stories. huggingface.co/datasets/kerinin/hackernews-stories.

Bài viết liên quan