A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional Generation
Shashi Narayan, Gonçalo Simões, Yao Zhao, Joshua Maynez, Dipanjan Das, Michael Collins, Mirella Lapata
摘要
We propose Composition Sampling, a simple but effective method to generate diverse outputs for conditional generation of higher quality compared to previous stochastic decoding strategies. It builds on recently proposed plan-based neural generation models (Narayan et al., 2021) that are trained to first create a composition of the output and then generate by conditioning on it and the input. Our approach avoids text degeneration by first sampling a composition in the form of an entity chain and then using beam search to generate the best possible text grounded to this entity chain. Experiments on summarization (CNN/DailyMail and XSum) and question generation (SQuAD), using existing and newly proposed automatic metrics together with human-based evaluation, demonstrate that Composition Sampling is currently the best available decoding strategy for generating diverse meaningful outputs. 1 Holtzman et al. ( 2020 ) use the term 'degeneration' to describe automatically generated text that is generic, repetitive, and awkward for story continuation. These issues are less common in conditional generation. In our case, 'degenerate' refers to text unfaithful or inconsistent to the input.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Diverse Demonstrations Improve In-context Compositional GeneralizationItay Levy, Ben Bogin, Jonathan BerantACL 2023 · 被引用 51 次
- Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought MethodYiming Wang, Zhuosheng Zhang, Rui WangACL 2023 · 被引用 39 次
- Distilling Script Knowledge from Large Language Models for Constrained Language PlanningSiyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge 等ACL 2023 · 被引用 14 次
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityAdithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita SarawagiEMNLP 2023 · 被引用 13 次
- Towards Summary Candidates FusionMathieu Ravaut, Shafiq R. Joty, Nancy F. ChenEMNLP 2022 · 被引用 10 次
它引用的顶会 Paper10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
相关 Paper
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- Conditional Poisson Stochastic BeamsClara Meister, Afra Amini, Tim Vieira, Ryan CotterellEMNLP 2021
- KCS: Diversify Multi-hop Question Generation with Knowledge Composition SamplingYangfan Wang, Jie Liu, Chen Tang, Lian Yan 等EMNLP 2025 · 被引用 1 次
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 被引用 21 次
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingWeizhen Qi, Yeyun Gong, Jian Jiao, Yu Yan 等ICML 2021 · 被引用 54 次
