Threads of Subtlety: Detecting Machine-Generated Texts Through Discourse Motifs
Zae Myung Kim, Kwang Hee Lee, Preston Zhu, Vipul Raheja, Dongyeop Kang
Abstract
With the advent of large language models (LLM), the line between human-crafted and machine-generated texts has become increasingly blurred. This paper delves into the inquiry of identifying discernible and unique linguistic properties in texts that were written by humans, particularly uncovering the underlying discourse structures of texts beyond their surface structures. Introducing a novel methodology, we leverage hierarchical parse trees and recursive hypergraphs to unveil distinctive discourse patterns in texts produced by both LLMs and humans. Empirical findings demonstrate that, although both LLMs and humans generate distinct discourse patterns influenced by specific domains, human-written texts exhibit more structural variability, reflecting the nuanced nature of human writing in different domains. Notably, incorporating hierarchical discourse features enhances binary classifiers' overall performance in distinguishing between human-written and machine-generated texts, even on out-of-distribution and paraphrased samples. This underscores the significance of incorporating hierarchical discourse features in the analysis of text patterns. The code and dataset are available at https://github.com/ minnesotanlp/threads-of-subtlety .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d83e403-b14d-4e16-a335-e4d62a1ec8dfCited by top-tier papers6
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba et al.ACL 2026 · 7 citations
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 3 citations
- Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text DetectionYang Li, Qiang Sheng, Zhengjia Wang, Yehan Yang et al.ACL 2026
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and AttributionLucio La Cava, Andrea TagarelliEMNLP 2025
- OSTAR: Optimized Statistical Text-classifier with Adversarial ResistanceYuhan Yao, Feifei Kou, Lei Shi, Xiao Yang et al.NeurIPS 2025
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
Related papers
- Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti et al.ACL 2026 · 3 citations
- Multi-level Style Preference Optimization: An Adaptive Detection Framework for Human-Machine Hybrid TextZehao Wang, Lianwei Wu, Wenbo An, Hang Zhang et al.AAAI 2026
- Comparing LLM-generated and human-authored news text using formal syntactic theoryOlga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-RodríguezACL 2025 · 8 citations
- Beyond Checkmate: Exploring the Creative Choke Points for AI Generated TextsNafis Irtiza Tripto, Saranya Venkatraman, Mahjabin Nahar, Dongwon LeeEMNLP 2025 · 1 citation
- PaLD: Detection of Text Partially Written by Large Language ModelsEric Lei, Hsiang Hsu, Chun-Fu ChenICLR 2025
