Home / Current Issue / Paper 1723524
Beyond Clean Benchmarks: A Reproducible Severity-Sweep Evaluation of AI-Generated Image Detector Robustness
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
AI-generated image detection technology is being adopted more frequently in journalism, content moderation, and digital-evidence verification, although the majority of AI detectors are tested on pristine, unaltered visuals and not on resized, recompressed, resampled, or altered formats one frequently encounters in the real world. Prior research, including the expansive multi-scenario RRDataset/RRBench (including seventeen detectors and more than ten vision-language models), has demonstrated that the performance of detectors deteriorates with image degradation; thus, this paper cannot claim to be the first work to show the robustness of detectors. Rather, it presents a concise, replicable evaluation of the two public detectors that function using differing architectures. — CNNSpot (a CNN created using ProGAN) and UnivFD (a frozen-CLIP semantic detector) were evaluated in a set of 2,000 images harmonized to derive a GenImage benchmark (Stable Diffusion v1.5, ADM, Midjourney, BigGAN). Both models were analyzed as frozen under a variety of conditions including JPEG compression of varying severity, image resizing and other manipulations. The results indicated that UnivFD was more accurate than CNNSpot in detection of original images (AUROC 0.741 vs. 0.662); this trend remains the same regardless of the 21 transformations. Gaussian blur was the only transformation that caused consistent monotonic drop in the performance. The JPEG compression, the resizing, and the noise led to minor changes that were non-monotonic. In addition, the cropping showed an improvement in AUROC scores for both detectors. Notably, however, no transformation crossed the threshold for practical significance set at ten. The compound chain of transformations did not hurt performance; where significant, compounding was mostly linked to the equal or better AUROC than the working of the worst.
Keywords
Detection of AI-generated images, image forensics, assessment of robustness, JPEG compression, distribution shift, Responsible AI, reproducibility.
References
[1] S.-Y. Wang, O. Wang, R. Zhang, A. Owens, and A. A. Efros, “CNN-generated images are surprisingly easy to spot... for now,” in Proc. IEEE/CVF CVPR, 2020. IEEE
[2] U. Ojha, Y. Li, and Y. J. Lee, “Towards universal fake image detectors that generalize across generative models,” in Proc. IEEE/CVF CVPR, 2023. IEEE
[3] Z. Wang et al., “DIRE for diffusion-generated image detection,” in Proc. IEEE/CVF ICCV, 2023. IEEE
[4] TrueMedia.org, “Distil-DIRE: a lightweight distilled variant of DIRE for AI-generated image detection,” official implementation. GitHub
[5] S. Yan et al., “AIDE: a sanity check for AI-generated image detection,” in Proc. ICLR, 2025. ICLR
[6] M. Zhu et al., “GenImage: a million-scale benchmark for detecting AI-generated image,” in Proc. NeurIPS, 2023. NeurIPS
[7] C. Li et al., “Bridging the gap between ideal and real-world evaluation: benchmarking AI-generated image detection in challenging scenarios,” in Proc. IEEE/CVF ICCV, 2025. CVF
[8] K. Li et al., “Detecting compressed AI-generated images via phase spectrum robustness,” in Proc. IEEE/CVF CVPR, 2026. CVF
[9] S. Choi, H. Lee, and M. Lee, “Training-free detection of AI-generated images via cropping robustness,” in Proc. NeurIPS, 2025. arXiv
[10] L. Pellegrini et al., “AI-GenBench: a new ongoing benchmark for AI-generated image detection,” in Proc. IJCNN, 2025. IEEE
[11] E. R. DeLong, D. M. DeLong, and D. L. Clarke-Pearson, “Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach,” Biometrics, vol. 44, no. 3, pp. 837-845, 1988. JSTOR
[12] Y. Benjamini and Y. Hochberg, “Controlling the false discovery rate: a practical and powerful approach to multiple testing,” J. R. Stat. Soc. Series B, vol. 57, no. 1, pp. 289-300, 1995. Wiley
How to cite this paper
@article{1723524,
author = {Aaranya Singh, Prakash Chand Jain, Ayush Singh},
title = {Beyond Clean Benchmarks: A Reproducible Severity-Sweep Evaluation of AI-Generated Image Detector Robustness},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {3},
pages = {3967-3974},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1723524.pdf},
abstract = {AI-generated image detection technology is being adopted more frequently in journalism, content moderation, and digital-evidence verification, although the majority of AI detectors are tested on pristine, unaltered visuals and not on resized, recompressed, resampled, or altered formats one frequently encounters in the real world. Prior research, including the expansive multi-scenario RRDataset/RRBench (including seventeen detectors and more than ten vision-language models), has demonstrated that the performance of detectors deteriorates with image degradation; thus, this paper cannot claim to be the first work to show the robustness of detectors. Rather, it presents a concise, replicable evaluation of the two public detectors that function using differing architectures. — CNNSpot (a CNN created using ProGAN) and UnivFD (a frozen-CLIP semantic detector) were evaluated in a set of 2,000 images harmonized to derive a GenImage benchmark (Stable Diffusion v1.5, ADM, Midjourney, BigGAN). Both models were analyzed as frozen under a variety of conditions including JPEG compression of varying severity, image resizing and other manipulations. The results indicated that UnivFD was more accurate than CNNSpot in detection of original images (AUROC 0.741 vs. 0.662); this trend remains the same regardless of the 21 transformations. Gaussian blur was the only transformation that caused consistent monotonic drop in the performance. The JPEG compression, the resizing, and the noise led to minor changes that were non-monotonic. In addition, the cropping showed an improvement in AUROC scores for both detectors. Notably, however, no transformation crossed the threshold for practical significance set at ten. The compound chain of transformations did not hurt performance; where significant, compounding was mostly linked to the equal or better AUROC than the working of the worst.},
keywords = {Detection of AI-generated images, image forensics, assessment of robustness, JPEG compression, distribution shift, Responsible AI, reproducibility.},
month = {September},
}