Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language Models
Sergio E. Zanotto, Segun Aroyehun
Abstract
The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. While recent research has primarily focused on using LLMs to classify text as either human-written or machine-generated texts, our study focuses on characterizing these texts using a set of linguistic features across different linguistic levels such as morphology, syntax, and semantics. We select a dataset of human-written and machine-generated texts spanning 8 domains and produced by 11 different LLMs. We calculate different linguistic features such as dependency length and emotionality, and we use them for characterizing human-written and machine-generated texts along with different sampling strategies, repetition controls, and model release dates. Our statistical analysis reveals that human-written texts tend to exhibit simpler syntactic structures and more diverse semantic content. Furthermore, we calculate the variability of our set of features across models and domains. Both human-and machinegenerated texts show stylistic diversity across domains, with human-written texts displaying greater variation in our features. Finally, we apply style embeddings to further test variability among human-written and machine-generated texts. Notably, newer models output text that is similarly variable, pointing to a homogenization of machine-generated texts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 110 citations
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 21 citations
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu et al.ACL 2024 · 18 citations
Related papers
- Threads of Subtlety: Detecting Machine-Generated Texts Through Discourse MotifsZae Myung Kim, Kwang Hee Lee, Preston Zhu, Vipul Raheja et al.ACL 2024 · 1 citation
- Multi-level Style Preference Optimization: An Adaptive Detection Framework for Human-Machine Hybrid TextZehao Wang, Lianwei Wu, Wenbo An, Hang Zhang et al.AAAI 2026
- Comparing LLM-generated and human-authored news text using formal syntactic theoryOlga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-RodríguezACL 2025 · 8 citations
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
- An Empirical Analysis of the Writing Styles of Persona-Assigned LLMsManuj Malik, Jing Jiang, Kian Ming A. ChaiEMNLP 2024 · 2 citations
