Hybrid Deep Learning Model for Fake Image Detection with Advanced Face Object Segmentation Method
Main Article Content
Abstract
The fast development of deepfake technologies for image generation produces more and more realistic manipulated facial images that are harder to distinguish from the real content. In this paper, we propose a novel Hybrid CNN–Vision Transformer (HybCNNViT) framework for robust deepfake image detection by combining discriminative and generative learning mechanisms. The proposed method presents a new dual-stream architecture that separately encodes inner facial components (eyes, nose and mouth) and outer facial regions with Variational Autoencoder (VAE) and Autoencoder (AE) modules respectively. Also, a hybrid CNN-ViT and CNN-Swin Transformer architecture is used to learn local visual artefacts and global contextual dependencies. In addition, a dedicated preprocessing scheme such as background removal and inner-outer facial object segmentation is further integrated to improve feature relevance and reduce computational complexity. The experimental results on real and fake face datasets under easy, medium and hard manipulation scenarios show the effectiveness of the proposed framework. The proposed Inner–Outer Segment Face Object + HybCNNViT model attains superior accuracy of 86%–88% and outperforms the existing CNN–ViT, ViViT, and CNN–ViTEfficientNet approaches. The results indicate that the combination of facial-region segmentation, latent feature learning and hybrid Transformer architectures enhances the robustness and generalisation capability of deepfake detection significantly.
