Detecting Deepfake and Synthetic Digital Content with Constrained Residual and Spectral Evidence
Main Article Content
Abstract
A photograph is no longer self-authenticating. Adversarial and diffusion generators produce faces that survive close human inspection, and identity transfer tools place a fabricated face inside an otherwise ordinary picture. Classifiers trained on visible appearance track this moving target poorly, because appearance is exactly what a generator is optimised to get right. This article describes CoSpec-Net, a detector that supplements appearance with two forms of evidence a generator does not explicitly optimise. A constrained prediction error filter, whose central tap is pinned and whose surrounding taps are renormalised at every forward pass, isolates the local departure of a pixel from its neighbourhood. An azimuthally averaged Fourier profile compresses the whole image into a radial description of how energy is distributed across spatial frequency, where resampling inside a generator leaves a characteristic tail. Rather than concatenating the three descriptions, CoSpec-Net lets them interrogate one another: appearance tokens query the forensic tokens and forensic tokens query the appearance tokens, for two exchange rounds, before a pooled decision is taken under a focal objective. Measured on 80,905 face images and on a separate corpus that distinguishes genuine, deepfake and wholly synthetic pictures, the detector attains 98.05 percent and 99.60 percent accuracy. Four published detectors report weaker figures on comparable material.
