Detection and Measurement of Syntactic Templates in Generated Text
Chantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. Wallace
Abstract
The diversity of text can be measured beyond word-level features, however existing diversity evaluation focuses primarily on word-level features. Here we propose a method for evaluating diversity over syntactic features to characterize general repetition in models, beyond frequent n-grams. Specifically, we define syntactic templates (e.g., strings comprising parts-of-speech) and show that models tend to produce templated text in downstream tasks at a higher rate than what is found in human-reference texts. We find that most (76%) templates in modelgenerated text can be found in pre-training data (compared to only 35% of human-authored text), and are not overwritten during fine-tuning or alignment processes such as RLHF. The connection between templates in generated text and the pre-training data allows us to analyze syntactic templates in models where we do not have the pre-training data. We also find that templates as features are able to differentiate between models, tasks, and domains, and are useful for qualitatively evaluating common model constructions. Finally, we demonstrate the use of templates as a useful tool for analyzing style memorization of training data in LLMs 1 . 1 https://cshaib.github.io/syntactic_templates/ … Mistral-7B VBZ DT JJ CC JJ 80/500 NN IN NN CC NN 70/500
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
- Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream TasksMiaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu et al.ACL 2026 · 16 citations
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 14 citations
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language ModelsChantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh GhassemiNeurIPS 2025 · 8 citations
- Comparing LLM-generated and human-authored news text using formal syntactic theoryOlga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-RodríguezACL 2025 · 8 citations
Builds on11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina et al.ICLR 2024 · 332 citations
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
Related papers
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 3 citations
- Repeated Sequences Reveal Gaps between Large Language Models and Natural LanguageKumiko Tanaka-IshiiACL 2026 · 1 citation
- Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language ModelsSamuel Paech, Allen Roush, Judah Goldfeder, Ravid Shwartz-ZivICLR 2026 · 2 citations
- Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition ProbabilityChristopher Nassif, Joshua CooperICML 2026
- RepEval: Effective Text Evaluation with LLM RepresentationShuqian Sheng, Yi Xu, Tianhang Zhang, Zanwei Shen et al.EMNLP 2024 · 5 citations
