Conditional set generation using Seq2seq models
Aman Madaan, Dheeraj Rajagopal, Niket Tandon, Yiming Yang, Antoine Bosselut
Abstract
Conditional set generation learns a mapping from an input sequence of tokens to a set. Several NLP tasks, such as entity typing and dialogue emotion tagging, are instances of set generation. Seq2Seq models are a popular choice to model set generation but they treat a set as a sequence and do not fully leverage its key properties, namely order-invariance and cardinality. We propose a novel algorithm for effectively sampling informative orders over the combinatorial space of label orders. Further, we jointly model the set cardinality and output by listing the set size as the first element and taking advantage of the autoregressive factorization used by Seq2Seq models. Our method is a model-independent data augmentation approach that endows any Seq2Seq model with the signals of order-invariance and cardinality. Training a Seq2Seq model on this new augmented data (without any additional annotations), gets an average relative improvement of 20% for four benchmarks datasets across models spanning from BART-base, T5-11B, and GPT-3. We will release all code and data upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1de96786-f613-40f5-a32e-0ea7bd3e6686Cited by top-tier papers3
- Enhancing Multi-Label Classification via Dynamic Label-Order LearningJiangnan Li, Yice Zhang, Shiwei Chen, Ruifeng XuAAAI 2024 · 4 citations
- To Code or Not To Code? Exploring Impact of Code in Pre-trainingViraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot et al.ICLR 2025 · 3 citations
- A Branching Decoder for Set GenerationZixian Huang, Gengyang Xiao, Yu Gu, Gong ChengICLR 2024 · 2 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
Related papers
- Target-Side Input Augmentation for Sequence to Sequence GenerationShufang Xie, Ang Lv, Yingce Xia, Lijun Wu et al.ICLR 2022 · 16 citations
- Precisely the Point: Adversarial Augmentations for Faithful and Informative Text GenerationWenhao Wu, Wei Li, Jiachen Liu, Xinyan Xiao et al.EMNLP 2022 · 4 citations
- Order-Agnostic Data Augmentation for Few-Shot Named Entity RecognitionHuiming Wang, Liying Cheng, Wenxuan Zhang, De Wen Soh et al.ACL 2024
- One2Set: Generating Diverse Keyphrases as a SetJiacheng Ye, Tao Gui, Yichao Luo, Yige Xu et al.ACL 2021
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang et al.AAAI 2024 · 30 citations
