Bài viết
Đánh giá hiệu quả các kĩ thuật phát hiện tin nhắn rác tiếng Việt
- Vu Minh Tuan (VN)
- Nguyen Xuan Thang (VN)
- Tran Quang Anh (VN)
Tóm tắt
Tóm tắt— Bài báo sử dụng cả mô hình học máy truyền thống và mô hình học sâu để phát hiện tin nhắn rác SMS tiếng Việt nhằm đánh giá hiệu quả các mô hình trên các dạng tập dữ liệu biến thể tiếng Việt khác nhau. Nhóm tác giả đã thử nghiệm 5 thuật toán Support Vector Machine (SVM), Naive Bayes (NB), Random Forests (RF), Convolutional Neural Network (CNN) và Long Short-Term Memory (LSTM) trên bộ dữ liệu tiếng Việt có dấu, tiếng Việt không dấu và tập hỗn hợp. Các phát hiện cho thấy LSTM và CNN, được hỗ trợ kỹ thuật học chuyển đổi PhoBert, hiệu quả hơn các mô hình học máy truyền thống. Mô hình LSTM cho độ chính xác cao nhất là 97,77% khi thử nghiệm trên bộ dữ liệu tiếng Việt đầy đủ dấu. Tương tự, mô hình CNN kết hợp với PhoBert cho độ chính xác cao nhất là 95,56% khi xử lý bộ dữ liệu tiếng Việt không dấu.
Lượt tải theo tháng
Di chuột vào cột để xem số lượt tải.
Cách trích dẫn
Vu Minh Tuan, Nguyen Xuan Thang, Tran Quang Anh (2023). Đánh giá hiệu quả các kĩ thuật phát hiện tin nhắn rác tiếng Việt. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 1(18), 30-37. https://doi.org/10.54654/isj.v1i18.932
Tài liệu tham khảo
- 1.CTIA, “2021 Annual Survey HIGHLIGHTS,” 2021. [Online]. Available: https://www.ctia.org/news/2021-annual-survey-highlights.
- 2.Attentive, “2021 SMS Marketing Benchmarks Report,” 2021. [Online]. Available: https://www.attentivemobile.com/2021-sms-marketing-benchmarks-report. [Accessed 2022].
- 3.Shafi’I Muhammad Abdulhamid; Muhammad Shafie Abd Latiff; Haruna Chiroma; Oluwafemi Osho; Gaddafi Abdul-Salaam, “A Review on Mobile SMS Spam Filtering Techniques,” IEEE Access, vol. 5, pp. 15650 - 15666, 2017.
- 4.K. Yadav, S. K. Saha, P. Kumaraguru, and R. Kumra, “Take control of your smses: Designing an usable spam sms filtering system,” in 2012 IEEE 13th International Conference on Mobile Data Management, Bengaluru, India, 2012.
- 5.El-Alfy, E.-S.M. and AlHasan, A.A., “Spam filtering framework for multimodal mobile communication based on dendritic cell algorithm,” Future Generation Computer Systems, vol. 64, pp. 98-107, 2016.
- 6.A. Narayan and P. Saxena, “The curse of 140 characters: evaluating the efficacy of sms spam detection on android,” in Third ACM workshop on Security and privacy in smartphones & mobile devices, Berlin, Germany, 2013.
- 7.Milivoje Popovac, Mirjana Karanovic, Srdjan Sladojevic, Marko Arsenovic, Andras Anderla, “Convolutional Neural Network Based SMS Spam Detection,” in 2018 26th Telecommunications Forum (TELFOR), Belgrade, Serbia , 2018.
- 8.Gauri Jain, Manisha Sharma, Basant Agarwal , “Optimizing semantic LSTM for spam detection,” International Journal of Information Technology, vol. 11, pp. 239 - 250, 2019.
- 9.W. Gomaa, “The Impact of Deep Learning Techniques on SMS Spam Filtering,” International Journal of Advanced Computer Science and Applications, vol. 11, no. 1, pp. 544 - 549, 2020.
- 10.Aliaksandr Barushka, Petr Hajek, “Spam filtering using integrated distribution-based balancing approach and regularized deep neural networks,” Applied Intelligence , vol. 48, p. 3538–3556, 2018.
- 11.. Vu Minh Tuan, Dang Dinh Quan, Nguyen Thanh Ha, Tran Quang Anh, “Lọc tin nhắn rác với Spam-Assassin,” Journal of Science and Technology on Information and Communications, vol. 3, no. 4, pp. 34-41, 2017.
- 12.Vu Minh Tuan, Quang Anh Tran, Minh Quang Ha, Lam Bui Thu, “A Multi-objective Approach for Vietnamese Spam Detection,” in Knowledge and Systems Engineering 2013, Hanoi, 2014.
- 13.Thai Hoang Pham, Phuong Le Hong, “Content-based Approach for Vietnamese Spam SMS Filtering,” in The 20th International Conference on Asian Language , Taiwain, 2016.
- 14.R. Johnson, T. Zhang, “Supervised and semi-supervised text categorization using LSTM for region embeddings,” in The 33rd International Conference on Machine Learning, New York, 2016.
- 15.X. Zhang, J. Zhao, Y. LeCun, “Character-level convolutional networks for text classification,” in The 28th Advances in Neural Information Processing Systems, Quebec, 2015.
- 16.Kiem-Hieu Nguyen, Cheol-Young Ock, “Diacritics Restoration in Vietnamese: Letter Based vs. Syllable Based Model,” in PRICAI 2010: Trends in Artificial Intelligence, Berlin, Heidelberg, 2010.
- 17.Jakub Náplava, Milan Straka, Pavel Straňák, Jan Hajič, “Diacritics Restoration Using Neural Networks,” in the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, 2018.
- 18.Hilal Tekgöz; Halil İbrahim Çelenli; Sevinç İlhan Omurca, “Semantic Similarity Comparison of Word Representation Methods in the Field of Health,” in 2021 6th International Conference on Computer Science and Engineering (UBMK), Ankara, Turkey, 2021.
- 19.Dat Quoc Nguyen, Anh Tuan Nguyen, “PhoBERT: Pre-trained language models for Vietnamese,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020.
- 20.G. Forman, “BNS feature scaling: an improved representation over tf-idf for svm text classification,” in Proceedings of the 17th ACM conference on Information and knowledge management, Napa Valley California USA, 2008.
- 21.J.A.K. Suykens; J. Vandewalle , “Least Squares Support Vector Machine Classifiers,” Neural Processing Letters , vol. 9, pp. 293 - 300, 1999.
- 22.George H. John, Pat Langley, “Estimating Continuous Distributions in Bayesian Classifiers,” in Eleventh Conference on Uncertainty in Artificial Intelligence (UAI1995), Quebec, Canada, 1995.
- 23.L. Breiman, “Random Forests,” Machine Learning volume , vol. 45, no. 1, pp. 5-32, 2001.