Attend to the Active: Structure-Aware Dynamic Attention in LLMs for Compositional Instruction Following
Fangrui Lv, Yulei Qin, Ruixin Hong, Jian Liang, Jinyang Wu, Ke Li, Xing Sun, Changshui Zhang
摘要
Large language models (LLMs) have demonstrated strong instruction-following capabilities; however, they often struggle with compositional instructions that involve multiple interleaved yet logically independent sub-tasks. These sub-tasks are typically organized in mutually exclusive structures, such as branching, chaining, or paralleling, where only one sub-task should be active at each generation step, while the others remain dormant. Despite their inactivity, dormant sub-tasks can inadvertently attract the model's attention due to structural entanglement within the input context or intermediate representations, leading to interference that compromises output fidelity. To address this challenge, we propose ATA, a structure-aware dynamic attention mechanism grounded in compositional structures, which dynamically identifies the active sub-task during generation while suppressing attention to inactive ones. By precisely steering the model's focus, ATA mitigates interference and explicitly enhances model adherence to the active sub-task. Importantly, ATA operates within a single forward pass without requiring parameter updates. Extensive experiments show that ATA consistently enhances LLMs' instructionfollowing ability across various compositional structures, effectively mitigating attention distraction and demonstrating a strong generalization ability. * Corresponding author. the description should be in English Wrong Generation If the work contains any animal, the description should be in Chinese. Otherwise, the description should be in English.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu 等ACL 2023 · 被引用 249 次
相关 Paper
- Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingSenyu Han, Hongchuan Zeng, Kai Yu, Lu ChenICML 2025
- Don't Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language ModelsZitong Shi, Frank Wan, Haixin Wang, Ruoyan Li 等NeurIPS 2025 · 被引用 2 次
- AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLMLian Zhong, Ziqiang He, Jibin Zheng, Jin Li 等CVPR 2026 · 被引用 2 次
- -Attn: Decomposed Attention for Large Vision-and-Language ModelsChia-Wen Kuo, Sijie Zhu, Fan Chen, Xiaohui Shen 等ICCV 2025 · 被引用 1 次
- PADA-Coder: Improving Plan-Following Code Generation via Perturbation-Verified Attention Distillation and Dynamic AlignmentYihong Huang, KE QIN, Rongzheng Wang, Muquan Li 等ICML 2026
