Deepfake Text Detection: Limitations and Opportunities
Jiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman, Yoonjin Kim, Parantapa Bhattacharya, Mobin Javed, Bimal Viswanath
Abstract
Recent advances in generative models for language have enabled the creation of convincing synthetic text or deepfake text. Prior work has demonstrated the potential for misuse of deepfake text to mislead content consumers. Therefore, deepfake text detection, the task of discriminating between human and machine-generated text, is becoming increasingly critical. Several defenses have been proposed for deepfake text detection. However, we lack a thorough understanding of their real-world applicability. In this paper, we collect deepfake text from 4 online services powered by Transformer-based tools to evaluate the generalization ability of the defenses on content in the wild. We develop several low-cost adversarial attacks, and investigate the robustness of existing defenses against an adaptive attacker. We find that many defenses show significant degradation in performance under our evaluation scenarios compared to their original claimed performance. Our evaluation shows that tapping into the semantic information in the text content is a promising approach for improving the robustness and generalization performance of deepfake text detection schemes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c1775d0-daf9-429e-86a0-bfbaae8104e7Cited by top-tier papers16
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang et al.ICLR 2024 · 311 citations
- A Representative Study on Human Detection of Artificially Generated Media Across CountriesJoel Frank, Franziska Herbert, Jonas Ricker, Lea Schönherr et al.S&P 2024 · 43 citations
- Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingTong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu et al.AAAI 2025 · 21 citations
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu et al.ACL 2024 · 18 citations
- Children, Parents, and Misinformation on Social MediaFilipo Sharevski, Jennifer Vander LoopS&P 2024 · 9 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Leveraging Frequency Analysis for Deep Fake Image RecognitionJoel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer et al.ICML 2020 · 848 citations
Related papers
- An Analysis of Recent Advances in Deepfake Image Detection in an Evolving Threat LandscapeSifat Muhammad Abdullah, Aravind Cheruvu, Shravya Kanchi, Taejoong Chung et al.S&P 2024 · 42 citations
- One Detector to Rule Them All: Towards a General Deepfake Attack Detection FrameworkShahroz Tariq, Sangyup Lee, Simon S. WooWWW 2021 · 90 citations
- Deepfake Videos in the Wild: Analysis and DetectionJiameng Pu, Neal Mangaokar, Lauren Kelly, Parantapa Bhattacharya et al.WWW 2021 · 59 citations
- "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial AttacksEdoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg GrohACL 2022 · 43 citations
- Detect Any AI-Counterfeited Text ImageChenfan Qu, Yiwu Zhong, Xuekang Zhu, Junchi Li et al.CVPR 2026
