Long Text Generation by Modeling Sentence-Level and Discourse-Level Coherence
Jian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu, Wenbiao Ding, Minlie Huang
Abstract
Generating long and coherent text is an important but challenging task, particularly for open-ended language generation tasks such as story generation. Despite the success in modeling intra-sentence coherence, existing generation models (e.g., BART) still struggle to maintain a coherent event sequence throughout the generated text. We conjecture that this is because of the difficulty for the decoder to capture the high-level semantics and discourse structures in the context beyond token-level co-occurrence. In this paper, we propose a long text generation model, which can represent the prefix sentences at sentence level and discourse level in the decoding process. To this end, we propose two pretraining objectives to learn the representations by predicting inter-sentence semantic similarity and distinguishing between normal and shuffled sentence orders. Extensive experiments show that our model can generate more coherent texts than state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang et al.NeurIPS 2024 · 151 citations
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationJin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai et al.NeurIPS 2022 · 135 citations
- HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event SequencesSiqiao Xue, Xiaoming Shi, James Y. Zhang, Hongyuan MeiNeurIPS 2022 · 65 citations
- Understanding In-Context Learning from RepetitionsJianhao Yan, Jin Xu, Chiyu Song, Chenming Wu et al.ICLR 2024 · 34 citations
- Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A SurveyJiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou et al.ACL 2024 · 15 citations
Builds on9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Content Planning for Neural Story Generation with Aristotelian RescoringSeraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph M. Weischedel, Nanyun PengEMNLP 2020 · 106 citations
- MEGATRON-CNTRL: Controllable Story Generation with External Knowledge Using Large-Scale Language ModelsPeng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri et al.EMNLP 2020 · 104 citations
Related papers
- DiscoDVT: Generating Long Text with Discourse-Aware Discrete Variational TransformerHaozhe Ji, Minlie HuangEMNLP 2021 · 17 citations
- DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank UtterancesXiaodong Gu, Kang Min Yoo, Jung-Woo HaAAAI 2021 · 83 citations
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 2 citations
- PAIR: Planning and Iterative Refinement in Pre-trained Transformers for Long Text GenerationXinyu Hua, Lu WangEMNLP 2020 · 44 citations
- StoryTrans: Non-Parallel Story Author-Style Transfer with Discourse Representations and Content EnhancingXuekai Zhu, Jian Guan, Minlie Huang, Juan LiuACL 2023 · 5 citations
