Attention Biasing and Context Augmentation for Zero-Shot Control of Encoder-Decoder Transformers for Natural Language Generation
Devamanyu Hazarika, Mahdi Namazifar, Dilek Hakkani-Tür
摘要
Controlling neural network-based models for natural language generation (NLG) to realize desirable attributes in the generated outputs has broad applications in numerous areas such as machine translation, document summarization, and dialog systems. Approaches that enable such control in a zero-shot manner would be of great importance as, among other reasons, they remove the need for additional annotated data and training. In this work, we propose novel approaches for controlling encoder-decoder transformer-based NLG models in zero shot. While zero-shot control has previously been observed in massive models (e.g., GPT3), our method enables such control for smaller models. This is done by applying two control knobs, attention biasing and context augmentation, to these models directly during decoding and without additional training or auxiliary models. These knobs control the generation process by directly manipulating trained NLG models (e.g., biasing cross-attention layers). We show that not only are these NLG models robust to such manipulations, but also their behavior could be controlled without an impact on their generation performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- Target-Side Input Augmentation for Sequence to Sequence GenerationShufang Xie, Ang Lv, Yingce Xia, Lijun Wu 等ICLR 2022 · 被引用 16 次
- Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention ModelsChungpa Lee, Jy-yong Sohn, Kangwook LeeICML 2026 · 被引用 1 次
- CoCon: A Self-Supervised Approach for Controlled Text GenerationAlvin Chan, Yew-Soon Ong, Bill Pung, Aston Zhang 等ICLR 2021 · 被引用 16 次
- Controllable Meaning Representation to Text Generation: Linearization and Data Augmentation StrategiesChris Kedzie, Kathleen R. McKeownEMNLP 2020 · 被引用 16 次
- ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue GenerationSeungmin Shin, Dooyoung Kim, Youngjoong KoEMNLP 2025
