"Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake Detectors
Kevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski, Allison Lu, Caroline Fedele, Magdalena Pasternak, Seth Layton, Kevin R. B. Butler, Carrie Gates, Patrick Traynor
Abstract
Audio deepfakes represent a rising threat to trust in our daily communications. In response to this, the research community has developed a wide array of detection techniques aimed at preventing such attacks from deceiving users. Unfortunately, the creation of these defenses has generally overlooked the most important element of the system - the user themselves. As such, it is not clear whether current mechanisms augment, hinder, or simply contradict human classification of deepfakes. In this paper, we perform the first large-scale user study on deepfake detection. We recruit over 1,200 users and present them with samples from the three most widely-cited deepfake datasets. We then quantitatively compare performance and qualitatively conduct thematic analysis to motivate and understand the reasoning behind user decisions and differences from machine classifications. Our results show that users correctly classify human audio at significantly higher rates than machine learning models, and rely on linguistic features and intuition when performing classification. However, users are also regularly misled by pre-conceptions about the capabilities of generated audio (e.g., that accents and background sounds are indicative of humans). Finally, machine learning models suffer from significantly higher false positive rates, and experience false negatives that humans correctly classify when issues of quality or robotic characteristics are reported. By analyzing user behavior across multiple deepfake datasets, our study demonstrates the need to more tightly compare user and machine learning performance, and to target the latter towards areas where humans are less likely to successfully identify threats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bd1614c-7205-41d1-af8d-2b7357d671d4Cited by top-tier papers2
- Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos DetectionChen Chen, Dion GohCHI 2026 · 2 citations
- Playing the Imitation Game: How Perceived Generated Content Shapes Player ExperienceMahsa Bazzaz, Seth CooperCHI 2026
Builds on12
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 187 citations
- SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification SystemsHadi Abdullah, Kevin Warren, Vincent Bindschaedler, Nicolas Papernot et al.S&P 2021 · 145 citations
- DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake VoicesRun Wang, Felix Juefei-Xu, Yihao Huang, Qing Guo et al.ACM MM 2020 · 124 citations
- Secure Your Voice: An Oral Airflow-Based Continuous Liveness Detection for Voice AssistantsYao Wang, Wandong Cai, Tao Gu, Wei Shao et al.UbiComp 2020 · 47 citations
Related papers
- Who Are You (I Really Wanna Know)? Detecting Audio DeepFakes Through Vocal Tract ReconstructionLogan Blue, Kevin Warren, Hadi Abdullah, Cassidy Gibson et al.USENIX Security 2022
- DeepPhish: Understanding User Trust Towards Artificially Generated Profiles in Online Social NetworksJaron Mink, Licheng Luo, Natã M. Barbosa, Olivia Figueira et al.USENIX Security 2022
- A Representative Study on Human Detection of Artificially Generated Media Across CountriesJoel Frank, Franziska Herbert, Jonas Ricker, Lea Schönherr et al.S&P 2024 · 43 citations
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang et al.WWW 2023 · 34 citations
- What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech DetectionBinh Nguyen, Shuju Shi, Ryan Ofman, Thai LeEMNLP 2025 · 1 citation
