USENIX Security2024Top-tier venue
SoK: The Good, The Bad, and The Unbalanced: Measuring Structural Limitations of Deepfake Media Datasets
Seth Layton, Tyler Tucker, Daniel Olszewski, Kevin Warren, Kevin R. B. Butler, Patrick Traynor
Abstract
Deepfake media represents an important and growing threat not only to computing systems but to society at large. Datasets of image, video, and voice deepfakes are being created to assist researchers in building strong defenses against these emerging threats. However, despite the growing number of datasets and the relative diversity of their samples, little guidance exists to help researchers select datasets and then meaningfully contrast their results against prior efforts. To assist in this process, this paper presents the first systematization of deepfake media. Using traditional anomaly detection datasets as a baseline, we characterize the metrics, generation techniques, and class distributions of existing datasets. Through this process, we discover significant problems impacting the comparability of systems using these datasets, including unaccounted-for heavy class imbalance and reliance upon limited metrics. These observations have a potentially profound impact should such systems be transitioned to practice -as an example, we demonstrate that the widely-viewed best detector applied to a typical call center scenario would result in only 1 out of 333 flagged results being a true positive. To improve reproducibility and future comparisons, we provide a template for reporting results in this space and advocate for the release of model score files such that a wider range of statistics can easily be found and/or calculated. Through this, and our recommendations for improving dataset construction, we provide important steps to move this community forward.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbc5a2fe-131d-44c7-925b-18903b7a17aeCited by top-tier papers5
- Collab: Fostering Critical Identification of Deepfake Videos on Social Media via Synergistic AnnotationShuning Zhang, Linzhi Wang, Shixuan Li, Yuanyuan Wu et al.CHI 2026 · 1 citation
- Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear ImplantMagdalena Pasternak, Kevin Warren, Daniel Olszewski, Susan Nittrouer et al.NDSS 2025
- Data to Infinity and Beyond: Examining Data Sharing and Reuse Practices in the Computer Security CommunityAnna Crowder, Allison Lu, Kevin Childs, Carson Stillman et al.S&P 2025
- "Helps me Take the Post With a Grain of Salt: " Soft Moderation Effects on Accuracy Perceptions and Sharing Intentions of Inauthentic Political Content on XFilipo Sharevski, Verena Distler, Florian AltUSENIX Security 2025
- SoK: Towards a Unified Approach to Applied Replicability for Computer SecurityDaniel Olszewski, Tyler Tucker, Kevin R. B. Butler, Patrick TraynorUSENIX Security 2025
Builds on10
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionBojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma et al.ACM MM 2020 · 443 citations
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder et al.USENIX Security 2019 · 441 citations
- KoDF: A Large-scale Korean DeepFake Detection DatasetPatrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park et al.ICCV 2021 · 154 citations
- Robust Performance Metrics for Authentication SystemsShridatt Sugrim, Can Liu, Meghan McLean, Janne LindqvistNDSS 2019 · 46 citations
Related papers
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat et al.ACM MM 2024 · 51 citations
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetKartik Thakral, Rishabh Ranjan, Akanksha Singh, Akshat Jain et al.ICLR 2025
- Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised LearningStefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta OneataCVPR 2025
- DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake DetectionChaoyu Gong, Han Zhang, Siqiang LuoCVPR 2026
