Few-Shot Detection of Machine-Generated Text using Style Representations
Rafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen, Marcus Bishop, Nicholas Andrews
摘要
The advent of instruction-tuned language models that convincingly mimic human writing poses a significant risk of abuse. However, such abuse may be counteracted with the ability to detect whether a piece of text was composed by a language model rather than a human author. Some previous approaches to this problem have relied on supervised methods by training on corpora of confirmed human- and machine- written documents. Unfortunately, model under-specification poses an unavoidable challenge for neural network-based detectors, making them brittle in the face of data shifts, such as the release of newer language models producing still more fluent text than the models used to train the detectors. Other approaches require access to the models that may have generated a document in question, which is often impractical. In light of these challenges, we pursue a fundamentally different approach not relying on samples from language models of concern at training time. Instead, we propose to leverage representations of writing style estimated from human-authored text. Indeed, we find that features effective at distinguishing among human authors are also effective at distinguishing human from machine authors, including state-of-the-art large language models like Llama-2, ChatGPT, and GPT-4. Furthermore, given a handful of examples composed by each of several specific language models of interest, our approach affords the ability to predict which model generated a given document. The code and data to reproduce our experiments are available at https://github.com/LLNL/LUAR/tree/main/fewshot_iclr2024.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive LearningXun Guo, Yongxin He, Shan Zhang, Ting Zhang 等NeurIPS 2024 · 被引用 100 次
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu 等ACL 2024 · 被引用 18 次
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation LearningYongxin He, Shan Zhang, Yixuan Cao, Lei Ma 等NeurIPS 2025 · 被引用 13 次
- Are AI-Generated Text Detectors Robust to Adversarial Perturbations?Guanhua Huang, Yuchen Zhang, Zhe Li, Yongjian You 等ACL 2024 · 被引用 7 次
- Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and DomainsJunghwan Kim, Haotian Zhang, David JurgensEMNLP 2025 · 被引用 3 次
它引用的顶会 Paper6
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- CATER: Intellectual Property Protection on Text Generation APIs via Conditional WatermarksXuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu 等NeurIPS 2022 · 被引用 106 次
相关 Paper
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi 等ACL 2024 · 被引用 44 次
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi 等ICML 2024 · 被引用 262 次
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold 等ICLR 2024 · 被引用 173 次
- Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition ProbabilityChristopher Nassif, Joshua CooperICML 2026
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution ConsistencyZhaoheng Huang, Yutao Zhu, Ji-Rong Wen, Zhicheng DouEMNLP 2025
