Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated Text
Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, Chris Callison-Burch
Abstract
As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior work on human detection of generated text focuses on the case where an entire passage is either human-written or machine-generated. In this paper, we study a more realistic setting where text begins as human-written and transitions to being generated by state-of-the-art neural language models. We show that, while annotators often struggle at this task, there is substantial variance in annotator skill and that given proper incentives, annotators can improve at this task over time. Furthermore, we conduct a detailed comparison study and analyze how a variety of variables (model size, decoding strategy, fine-tuning, prompt genre, etc.) affect human detection performance. Finally, we collect error annotations from our participants and use them to show that certain textual genres influence models to make different types of errors and that certain sentence-level features correlate highly with annotator selection. We release the RoFT dataset: a collection of over 21,000 human annotations paired with error classifications to encourage future work in human detection and evaluation of generated text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4865098a-16dc-40ca-9d25-7a35aab9266aCited by top-tier papers13
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and ModelsJiachen Fu, Chun-Le Guo, Chongyi LiACM MM 2025 · 4 citations
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 3 citations
- IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text DetectorZheng Chen, Yushi Feng, Jisheng Dang, Changyang He et al.NeurIPS 2025 · 2 citations
- Boosting the Uniqueness of Neural Networks Fingerprints with Informative TriggersZhuomeng Zhang, Fangqi Li, Hanyi Wang, Shi-Lin WangNeurIPS 2025 · 1 citation
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 21 citations
- Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos et al.EMNLP 2020 · 12 citations
Related papers
- Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti et al.ACL 2026 · 3 citations
- HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI CoauthoringZhixiong Su, Yichen Wang, Herun Wan, Zhaohan Zhang et al.ACL 2025 · 10 citations
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical GuaranteesHongyi Zhou, Jin Zhu, Pingfan Su, Kai Ye et al.NeurIPS 2025 · 23 citations
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu et al.ACL 2024 · 18 citations
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
