Language Model Detectors Are Easily Optimized Against
Charlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma, Christopher D. Manning, Chelsea Finn, Stefano Ermon
摘要
The fluency and general applicability of large language models (LLMs) has motivated significant interest in detecting whether a piece of text was written by a language model. While both academic and commercial detectors have been deployed in some settings, particularly education, other research has highlighted the fragility of these systems. In this paper, we demonstrate a data-efficient attack that fine-tunes language models to confuse existing detectors, leveraging recent developments in reinforcement learning of language models. We use the 'human-ness' score (often just a log probability) of various open-source and commercial detectors as a reward function for reinforcement learning, subject to a KL-divergence constraint that the resulting model does not differ significantly from the original. For a 7B parameter Llama-2 model, fine-tuning for under a day reduces the AUROC of the OpenAI RoBERTa-Large detector from 0.84 to 0.63, while perplexity on OpenWebText increases from 8.7 to only 9.0; with a larger perplexity budget, we can drive AUROC to 0.30 (worse than random). Similar to traditional adversarial attacks, we find that this increase in 'detector evasion' generalizes to other detectors not used during training. In light of our empirical results, we advise against continued reliance on LLM-generated text detectors. Models, datasets, and selected experiment code will be released at https://github.com/charlottttee/llm-detector-evasion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Attacks on Machine-Text Detectors Retain Stylistic FingerprintsRafael Rivera Soto, Barry Chen, Nicholas AndrewsICML 2026 · 被引用 1 次
- Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text DetectorsHao Fang, Jiawei Kong, Tianqu Zhuang, Yixiang Qiu 等EMNLP 2025
- Humanizing the Machine: Proxy Attacks to Mislead LLM DetectorsTianchun Wang, Yuanzhou Chen, Zichuan Liu, Zhanwen Chen 等ICLR 2025
- DALD: Improving Logits-based Detector without Logits from Black-box LLMsCong Zeng, Shengkun Tang, Xianjun Yang, Yuanzhou Chen 等NeurIPS 2024
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
相关 Paper
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha 等NeurIPS 2025 · 被引用 28 次
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang 等EMNLP 2023 · 被引用 21 次
- RAFT: Realistic Attacks to Fool Text DetectorsJames Wang, Ran Li, Junfeng Yang, Chengzhi MaoEMNLP 2024 · 被引用 2 次
- Learning to Rewrite: Generalized LLM-Generated Text DetectionWei Hao, Ran Li, Weiliang Zhao, Junfeng Yang 等ACL 2025
- Few-Shot Detection of Machine-Generated Text using Style RepresentationsRafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen 等ICLR 2024 · 被引用 49 次
