X-Deep Fake Net: An Explainable Transformer Framework for Deepfake Detection and Manipulation Localization

    DOI: https://doie.org/10.65985/JBSE.2026973194

    Rahul Burkul, Sandhya Waghere


    Keywords:

    Deepfake Detection, Swin Transformer, Explainable AI, Grad-CAM, Attention Rollout, Manipulation Localization, FaceForensics++, Vision Transformer, Image Forensics.


    Abstract:

    Deepfake technologies have become more advanced such that they cannot be differentiated from real media content anymore. Currently, most of the detection techniques concentrate on classifying whether the media is deepfake or not without having the interpretability and localization capability. In this paper, an explainable deepfake detection model called X DeepFakeNet++ is proposed where it incorporates a Swin transformer classifier, Grad-CAM, Attention Rollout, and manipulation localization component that was trained via automatically generated pseudo-masks from FaceForensics++ videos. Not only does the model detect deepfakes effectively (99.65%) but it also identifies the facial areas causing the detection as well as localized the manipulated areas.


    PDF

Indexed By