RAFT: Realistic Attacks to Fool Text Detectors
James Wang, Ran Li, Junfeng Yang, Chengzhi Mao
Abstract
Large language models (LLMs) have exhibited remarkable fluency across various tasks. However, their unethical applications, such as disseminating disinformation, have become a growing concern. Although recent works have proposed a number of LLM detection methods, their robustness and reliability remain unclear. In this paper, we present RAFT: a grammar error-free black-box attack against existing LLM detectors. In contrast to previous attacks for language models, our method exploits the transferability of LLM embeddings at the word-level while preserving the original text quality. We leverage an auxiliary embedding to greedily select candidate words to perturb against the target detector. Experiments reveal that our attack effectively compromises all detectors in the study across various domains by up to 99%, and are transferable across source models. Manual human evaluation studies show our attacks are realistic and indistinguishable from original human-written text. We also show that examples generated by RAFT can be used to train adversarially robust detectors. Our work shows that current LLM detectors are not adversarially robust, underscoring the urgent need for more resilient detection mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 569737f7-148c-446a-96c1-46782dbaa696Cited by top-tier papers4
- Frankentext: Stitching random text fragments into long-form narrativesChau Minh Pham, Jenna Russell, Dzung Pham, Mohit IyyerACL 2026 · 7 citations
- Attacking Misinformation Detection Using Adversarial Examples Generated by Language ModelsPiotr Przybyla, Euan McGill, Horacio SaggionEMNLP 2025 · 1 citation
- TempParaphraser: "Heating Up" Text to Evade AI-Text Detection through ParaphrasingJunjie Huang, Ruiquan Zhang, Jinsong Su, Yidong ChenEMNLP 2025
- Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text DetectorsHao Fang, Jiawei Kong, Tianqu Zhuang, Yixiang Qiu et al.EMNLP 2025
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
Related papers
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang et al.ICML 2026 · 4 citations
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha et al.NeurIPS 2025 · 28 citations
- Language Model Detectors Are Easily Optimized AgainstCharlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma et al.ICLR 2024 · 18 citations
- DA³: A Distribution-Aware Adversarial Attack against Language ModelsYibo Wang, Xiangjue Dong, James Caverlee, Philip S. YuEMNLP 2024 · 2 citations
- PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant AttacksZhenxin Ai, Haiyun HeICML 2026 · 4 citations
