Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
Yuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti, Zhuohan Xie, Minh Ngoc Ta, Jiahui Geng, Jinyan Su, Mervat Abassy, Saadeldine Eletter, Kareem Ashraf Elozeiri, Nurkhan Laiyk
Abstract
Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random guessing. To verify the generalizability of this finding across languages and domains, we perform an extensive case study to identify the upper bound of human detection accuracy. Across 16 datasets covering 9 languages and 9 domains, 19 annotators achieved an average detection accuracy of 87.6%, thus challenging previous conclusions. We find that major gaps between human and machine text lie in concreteness, cultural nuances, and diversity. Prompting by explicitly explaining the distinctions in the prompts can partially bridge the gaps in over 50% of the cases. However, we also find that humans do not always prefer human-written text, particularly when they cannot clearly identify its source. We release our dataset, the human labels, and the annotator metadata at https://github.com/ xnlp-lab/HumanEval-MGT . Setting ID Input Task Description Outputs Applicable Scenarios I. Single-Binary hwt or mgt Given a piece of text, identify whether it is written by human? A. Yes, human; B. No, machine parallel data is not necessary. II. Pair-Binary (hwt, mgt) or (mgt, hwt) Given a pair of (text1, text2), identify which one is humanwritten? Either text1 or text2 must be hwt, and another is mgt randomly sampled from mgti. A. text1; B. text2 parallel data is available. III. Triplet-Three-Class (hwt, mgt1, mgt2) Given a set of texts (text1, text2, text3), identify which one is human-written? One of the text1, text2 and text3 must be hwt, and others are mgt randomly sampled from mgti. A. text1; B. text2; C. text3 parallel human and multiple LLM generations are collected to make comparisons. IV. Pair-Four-Class (hwt, mgt) or (mgt, hwt) or (hwt, hwt) or (mgt, mgt) Given a pair of (text1, text2), identify which one is humanwritten? Both text1 and text2 can be hwt, and can be mgt. A. text1; B. text2; C. none; D. both parallel data is not necessary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f59b78e-3dd4-4911-afd2-8233c067c281Builds on8
- AI-Mediated Communication: Language Use and Interpersonal Effects in a Referential Communication TaskHannah Mieczkowski, Jeffrey T. Hancock, Mor Naaman, Malte F. Jung et al.CSCW 2021 · 83 citations
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
- Multi-Stage Pre-training for Automated Chinese Essay ScoringWei Song, Kai Zhang, Ruiji Fu, Lizhen Liu et al.EMNLP 2020 · 28 citations
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 21 citations
Related papers
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su et al.ACL 2024
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi et al.ICML 2024 · 262 citations
- Multi-level Style Preference Optimization: An Adaptive Detection Framework for Human-Machine Hybrid TextZehao Wang, Lianwei Wu, Wenbo An, Hang Zhang et al.AAAI 2026
- MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral AlignmentShengchao Liu, Xiaoming Liu, Chengzhengxu Li, Zhaohan Zhang et al.AAAI 2026 · 1 citation
