Bài viết

DSViT: Mô hình Transformer cải tiến cho phát hiện Deepfake

Từ khóaphát hiện deepfakeDSViTdeepfakedeepfake dựa trên không gian

Tóm tắt

Sự phát triển nhanh chóng của trí tuệ nhân tạo và các mô hình học sâu đã cho phép tạo ra các hình ảnh, video giả mạo siêu thực, đe dọa nghiêm trọng đến an toàn, an ninh thông tin. Việc phát hiện chính xác các hình ảnh, video giả mạo này là rất quan trọng để ngăn chặn lan truyền thông tin sai lệch và đảm bảo tính toàn vẹn của phương tiện kỹ thuật số. Hiện đã có nhiều các công trình nghiên cứu tiên tiến về phát hiện deepfake như ViT, CViT, song vẫn còn có những hạn chế nhất định, dẫn đến cần tiếp các công trình nghiên cứu để cải tiến thêm. Bài báo này đề xuất một mô hình dựa trên cải tiến từ mô hình CViT để tối ưu hóa cho việc phát hiện deepfake, gọi là DSViT (phát hiện deepfake với SC-based Convolutional Vision Transformer). Mô hình đề xuất sử dụng các khối Convolution và SCConvolution được sắp xếp một cách hợp lý kết hợp với kiến trúc ViT. Chúng tôi đã thử nghiệm trên bộ dữ liệu DFDC và so sánh kết quả với mô hình CViT để chứng minh hiệu quả của mô hình

Lượt tải theo tháng

0214110/2411/2412/2402/2503/2504/2505/2506/2507/2508/2509/2510/2511/2512/2501/2602/2603/2604/2605/26

Di chuột vào cột để xem số lượt tải.

Cách trích dẫn

Phạm Minh Thuấn, Bùi Thu Lâm, Phạm Duy Trung (2024). DSViT: Mô hình Transformer cải tiến cho phát hiện Deepfake. Tạp chí Khoa học và Công nghệ trong lĩnh vực An toàn thông tin, 2(22), 17-28. https://doi.org/10.54654/isj.v2i22.1055

Tài liệu tham khảo

  1. 1.F. Abbas and A. Taeihagh, “Unmasking deepfakes: A systematic review of deepfake detection and generation techniques using artificial intelligence,” Expert Systems With Applications, 2024: 124260.
  2. 2.A. Naitali, M. Ridouani, F. Salahdine, M. Kaabouch, “Deepfake attacks: Generation, detection, datasets, challenges, and research directions,” Computers, vol. 12, no. 10, pp. 216, Oct 2023.
  3. 3.X. Li, H. Zhou, and M. Zhao, “Transformer-based cascade networks with spatial and channel reconstruction convolution for deepfake detection,” Mathematical Biosciences and Engineering, vol. 21, no. 3, pp. 4142-4164, 2024.
  4. 4.D. Wodajo & S. Atnafu, “Deepfake Video Detection Using Convolutional Vision Transformer”, arXiv preprint arXiv:2102.11126, 2021.
  5. 5.F. Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul 2017.
  6. 6.H. H. Nguyen, N. T. Tieu & I. Echizen, “Capsule-Forensics: Using Capsule Networks to Detect Forged Images and Videos”, ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2307-2311, May 2019.
  7. 7.D. Afchar, V. Nozick, J. Yamagishi, and I. Echizen, “MesoNet: A Compact Facial Video Forgery Detection Network,” arXiv:1809.00888, Sep 2018.
  8. 8.M. Tan and Q. V. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” in International Conference on Machine Learning (ICML), Jun 2019.
  9. 9.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning Representations (ICLR), May 2021.
  10. 10.A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić & C. Schmid, “ViViT: A Video Vision Transformer”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6816-6826, 2021.
  11. 11.Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” in International Conference on Computer Vision (ICCV), Oct 2021.
  12. 12.W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), Feb 2022.
  13. 13.J. Li, Y. Wen, and L. He, “SCConv: Spatial and channel reconstruction convolution for feature redundancy,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  14. 14.Kagge (2020), Deepfake Detection Challenge. Accessed September 10, 2024, from: https://www.kaggle.com/c/deepfake-detection-challenge/data.

Bài viết liên quan