Coherence boosting: When your pretrained language model is not paying enough attention
Nikolay Malkin, Zhen Wang, Nebojsa Jojic
2022年份
45被引次数
15顶会引用
摘要
Long-range semantic coherence remains a challenge in automatic language generation and understanding. We demonstrate that large language models have insufficiently learned the effect of distant words on next-token prediction. We present coherence boosting, an inference procedure that increases a LM’s focus on a long context. We show the benefits of coherence boosting with pretrained models by distributional analyses of generated ordinary text and dialog responses. It is also found that coherence boosting with state-of-the-art models for various zero-shot NLP tasks yields performance gains with no additional training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt OptimizationXinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai 等ICLR 2024 · 被引用 226 次
- GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsZhen Wan, Fei Cheng, Zhuoyuan Mao, Qianying Liu 等EMNLP 2023 · 被引用 132 次
- Amortizing intractable inference in large language modelsEdward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar 等ICLR 2024 · 被引用 91 次
- Exposing Attention Glitches with Flip-Flop Language ModelingBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy 等NeurIPS 2023 · 被引用 90 次
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang 等NeurIPS 2025 · 被引用 33 次
它引用的顶会 Paper14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
相关 Paper
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- Provable Long-Range Benefits of Next-Token PredictionXinyuan Cao, Santosh S. VempalaSTOC 2026
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 被引用 15 次
- Long Text Generation by Modeling Sentence-Level and Discourse-Level CoherenceJian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu 等ACL 2021
