DOI: https://doie.org/10.65985/JBSE.2026973194
Rahul Burkul, Sandhya Waghere
Deepfake Detection, Swin Transformer, Explainable AI, Grad-CAM, Attention Rollout, Manipulation Localization, FaceForensics++, Vision Transformer, Image Forensics.
Deepfake technologies have become more advanced such that they cannot be differentiated from real media content anymore. Currently, most of the detection techniques concentrate on classifying whether the media is deepfake or not without having the interpretability and localization capability. In this paper, an explainable deepfake detection model called X DeepFakeNet++ is proposed where it incorporates a Swin transformer classifier, Grad-CAM, Attention Rollout, and manipulation localization component that was trained via automatically generated pseudo-masks from FaceForensics++ videos. Not only does the model detect deepfakes effectively (99.65%) but it also identifies the facial areas causing the detection as well as localized the manipulated areas.