Model Criticism for Long-Form Text Generation
Yuntian Deng, Volodymyr Kuleshov, Alexander M. Rush
Abstract
Language models have demonstrated the ability to generate highly fluent text; however, it remains unclear whether their output retains coherent high-level structure (e.g., story progression). Here, we propose to apply a statistical tool, model criticism in latent space, to evaluate the high-level structure of the generated text. Model criticism compares the distributions between real and generated data in a latent space obtained according to an assumptive generative process. Different generative processes identify specific failure modes of the underlying model. We perform experiments on three representative aspects of high-level discourse—coherence, coreference, and topicality—and find that transformer-based language models are able to capture topical structures but have a harder time maintaining structural coherence or modeling coreference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9266348-4bf8-4b51-b469-e1835639a815Cited by top-tier papers7
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 9 citations
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement LearningYuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee et al.ICLR 2026 · 9 citations
- BBScore: A Brownian Bridge Based Metric for Assessing Text CoherenceZhecheng Sheng, Tianhao Zhang, Chen Jiang, Dongyeop KangAAAI 2024 · 8 citations
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityMario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández et al.EMNLP 2023 · 5 citations
- Structure-Conditional Minimum Bayes Risk DecodingBryan Eikema, Anna Rutkiewicz, Mario GiulianelliEMNLP 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
Related papers
- DiscoDVT: Generating Long Text with Discourse-Aware Discrete Variational TransformerHaozhe Ji, Minlie HuangEMNLP 2021 · 17 citations
- Let Language Constrain Geometry: Vision–Language Models as Semantic and Spatial Critics for 3D GenerationWeimin Bai, Yubo Li, Weijian Luo, Zeqiang Lai et al.ICML 2026 · 1 citation
- Discourse Realization of Generics in Human and LLM-generated TextsSøren Kirkegaard Fomsgaard, Martial Pastor, Gaël Dias, Nelleke OostdijkACL 2026
- PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationZhe Hu, Hou Pong Chan, Jiachen Liu, Xinyan Xiao et al.ACL 2022
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
