Model Criticism for Long-Form Text Generation
Yuntian Deng, Volodymyr Kuleshov, Alexander M. Rush
摘要
Language models have demonstrated the ability to generate highly fluent text; however, it remains unclear whether their output retains coherent high-level structure (e.g., story progression). Here, we propose to apply a statistical tool, model criticism in latent space, to evaluate the high-level structure of the generated text. Model criticism compares the distributions between real and generated data in a latent space obtained according to an assumptive generative process. Different generative processes identify specific failure modes of the underlying model. We perform experiments on three representative aspects of high-level discourse—coherence, coreference, and topicality—and find that transformer-based language models are able to capture topical structures but have a harder time maintaining structural coherence or modeling coreference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 被引用 9 次
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement LearningYuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee 等ICLR 2026 · 被引用 9 次
- BBScore: A Brownian Bridge Based Metric for Assessing Text CoherenceZhecheng Sheng, Tianhao Zhang, Chen Jiang, Dongyeop KangAAAI 2024 · 被引用 8 次
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityMario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández 等EMNLP 2023 · 被引用 5 次
- Structure-Conditional Minimum Bayes Risk DecodingBryan Eikema, Anna Rutkiewicz, Mario GiulianelliEMNLP 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
相关 Paper
- DiscoDVT: Generating Long Text with Discourse-Aware Discrete Variational TransformerHaozhe Ji, Minlie HuangEMNLP 2021 · 被引用 17 次
- Let Language Constrain Geometry: Vision–Language Models as Semantic and Spatial Critics for 3D GenerationWeimin Bai, Yubo Li, Weijian Luo, Zeqiang Lai 等ICML 2026 · 被引用 1 次
- Discourse Realization of Generics in Human and LLM-generated TextsSøren Kirkegaard Fomsgaard, Martial Pastor, Gaël Dias, Nelleke OostdijkACL 2026
- PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationZhe Hu, Hou Pong Chan, Jiachen Liu, Xinyan Xiao 等ACL 2022
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
