Uncovering Human Traits in Determining Real and Spoofed Audio: Insights from Blind and Sighted Individuals
Chaeeun Han, Prasenjit Mitra, Syed Masum Billah
Abstract
This paper explores how blind and sighted individuals perceive real and spoofed audio, highlighting differences and similarities between the groups. Through two studies, we find that both groups focus on specific human traits in audio–such as accents, vocal inflections, breathing patterns, and emotions–to assess audio authenticity. We further reveal that humans, irrespective of visual ability, can still outperform current state-of-the-art machine learning models in discerning audio authenticity; however, the task proves psychologically demanding. Moreover, detection accuracy scores between blind and sighted individuals are comparable, but each group exhibits unique strengths: the sighted group excels at detecting deepfake-generated audio, while the blind group excels at detecting text-to-speech (TTS) generated audio. These findings not only deepen our understanding of machine-manipulated and neural-renderer audio but also have implications for developing countermeasures, such as perceptible watermarks and human-AI collaboration strategies for spoofing detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa4dbc0a-43ee-40dd-a3a0-d138de189030Cited by top-tier papers5
- Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal ModelsJingyi Xie, Rui Yu, He Zhang, Syed Masum Billah et al.CHI 2025 · 40 citations
- Characterizing Photorealism and Artifacts in Diffusion Model-Generated ImagesNegar Kamali, Karyn Nakamura, Aakriti Kumar, Angelos Chatzimparmpas et al.CHI 2025 · 23 citations
- SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content CreationStephen Brade, Sam Anderson, Rithesh Kumar, Zeyu Jin et al.CHI 2025 · 7 citations
- Effect of AI Performance, Risk Perception, and Trust on Human Dependence in Deepfake Detection AI SystemYingfan Zhou, Ester Chen, Manasa Pisipati, Aiping Xiong et al.CSCW 2025 · 2 citations
- Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear ImplantMagdalena Pasternak, Kevin Warren, Daniel Olszewski, Susan Nittrouer et al.NDSS 2025
Builds on2
Related papers
- Blind and Low-Vision Individuals' Detection of Audio DeepfakesFilipo Sharevski, Aziz Zeidieh, Jennifer Vander Loop, Peter JachimCCS 2024 · 5 citations
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional AnalysisKavita Kumari, Maryam Abbasihafshejani, Alessandro Pegoraro, Phillip Rieger et al.NDSS 2025
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera et al.ACM MM 2020 · 314 citations
