Language Model Detectors Are Easily Optimized Against
Charlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma, Christopher D. Manning, Chelsea Finn, Stefano Ermon
Abstract
The fluency and general applicability of large language models (LLMs) has motivated significant interest in detecting whether a piece of text was written by a language model. While both academic and commercial detectors have been deployed in some settings, particularly education, other research has highlighted the fragility of these systems. In this paper, we demonstrate a data-efficient attack that fine-tunes language models to confuse existing detectors, leveraging recent developments in reinforcement learning of language models. We use the 'human-ness' score (often just a log probability) of various open-source and commercial detectors as a reward function for reinforcement learning, subject to a KL-divergence constraint that the resulting model does not differ significantly from the original. For a 7B parameter Llama-2 model, fine-tuning for under a day reduces the AUROC of the OpenAI RoBERTa-Large detector from 0.84 to 0.63, while perplexity on OpenWebText increases from 8.7 to only 9.0; with a larger perplexity budget, we can drive AUROC to 0.30 (worse than random). Similar to traditional adversarial attacks, we find that this increase in 'detector evasion' generalizes to other detectors not used during training. In light of our empirical results, we advise against continued reliance on LLM-generated text detectors. Models, datasets, and selected experiment code will be released at https://github.com/charlottttee/llm-detector-evasion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e53f0a6-16c7-48d5-be81-9ea1584a8e37Cited by top-tier papers4
- Attacks on Machine-Text Detectors Retain Stylistic FingerprintsRafael Rivera Soto, Barry Chen, Nicholas AndrewsICML 2026 · 1 citation
- Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text DetectorsHao Fang, Jiawei Kong, Tianqu Zhuang, Yixiang Qiu et al.EMNLP 2025
- Humanizing the Machine: Proxy Attacks to Mislead LLM DetectorsTianchun Wang, Yuanzhou Chen, Zichuan Liu, Zhanwen Chen et al.ICLR 2025
- DALD: Improving Logits-based Detector without Logits from Black-box LLMsCong Zeng, Shengkun Tang, Xianjun Yang, Yuanzhou Chen et al.NeurIPS 2024
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha et al.NeurIPS 2025 · 28 citations
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang et al.EMNLP 2023 · 21 citations
- RAFT: Realistic Attacks to Fool Text DetectorsJames Wang, Ran Li, Junfeng Yang, Chengzhi MaoEMNLP 2024 · 2 citations
- Learning to Rewrite: Generalized LLM-Generated Text DetectionWei Hao, Ran Li, Weiliang Zhao, Junfeng Yang et al.ACL 2025
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen et al.ICLR 2024 · 49 citations
