Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated Text
Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, Chris Callison-Burch
摘要
As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior work on human detection of generated text focuses on the case where an entire passage is either human-written or machine-generated. In this paper, we study a more realistic setting where text begins as human-written and transitions to being generated by state-of-the-art neural language models. We show that, while annotators often struggle at this task, there is substantial variance in annotator skill and that given proper incentives, annotators can improve at this task over time. Furthermore, we conduct a detailed comparison study and analyze how a variety of variables (model size, decoding strategy, fine-tuning, prompt genre, etc.) affect human detection performance. Finally, we collect error annotations from our participants and use them to show that certain textual genres influence models to make different types of errors and that certain sentence-level features correlate highly with annotator selection. We release the RoFT dataset: a collection of over 21,000 human annotations paired with error classifications to encourage future work in human detection and evaluation of generated text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 被引用 39 次
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and ModelsJiachen Fu, Chun-Le Guo, Chongyi LiACM MM 2025 · 被引用 4 次
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 被引用 3 次
- IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text DetectorZheng Chen, Yushi Feng, Jisheng Dang, Changyang He 等NeurIPS 2025 · 被引用 2 次
- Boosting the Uniqueness of Neural Networks Fingerprints with Informative TriggersZhuomeng Zhang, Fangqi Li, Hanyi Wang, Shi-Lin WangNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 被引用 21 次
- Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos 等EMNLP 2020 · 被引用 12 次
相关 Paper
- Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti 等ACL 2026 · 被引用 3 次
- HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI CoauthoringZhixiong Su, Yichen Wang, Herun Wan, Zhaohan Zhang 等ACL 2025 · 被引用 10 次
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical GuaranteesHongyi Zhou, Jin Zhu, Pingfan Su, Kai Ye 等NeurIPS 2025 · 被引用 23 次
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu 等ACL 2024 · 被引用 18 次
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi 等ACL 2024 · 被引用 44 次
