LOGO - Long cOntext aliGnment via efficient preference Optimization
Zecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu, Min Zhang
摘要
Long-context models (LCMs) have shown great potential in processing long sequences, with research showing they can accurately locate tokenlevel salient information. Yet, the generation performance of these LCMs is far from satisfactory and might result in misaligned responses, such as hallucinations. To enhance the generation capability, existing works have investigated the effects of data size and quality for both pre-training and instruction tuning stages. Though achieving meaningful improvement, previous methods fall short in either effectiveness or efficiency. In this paper, we introduce LOGO, an efficient and effective training strategy that first introduces preference optimization for long-context alignment. LOGO consists of a reference-free preference optimization strategy and a corresponding efficient data synthesis process. By training with only 0.3B data on a single 8×A800 GPU machine for 16 hours, LOGO allows the Llama-3-8B-Instruct-80K model to achieve comparable performance with GPT-4 in real-world long-context tasks while preserving the model's original capabilities on other tasks, e.g., language modeling and MMLU. Besides, LOGO can also scale the models' context window size while enhancing their performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference OptimizationHuashan Sun, Shengyi Liao, Yansen Han, Yu Bai 等ICLR 2026 · 被引用 9 次
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition FrameworkZach Xu, Shang Zhu, Jue Wang, Junlin Wang 等ICLR 2026 · 被引用 9 次
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context InstructionsChaochen Gao, Xing Wu, Zijia Lin, Debing Zhang 等NeurIPS 2025 · 被引用 8 次
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference OptimizationShaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu 等ACL 2026 · 被引用 7 次
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
相关 Paper
- LongRecipe: Recipe for Efficient Long Context Generalization in Large Language ModelsZhiyuan Hu, Yuliang Liu, Jinman Zhao, Suyuchen Wang 等ACL 2025
- LIONs: An Empirically Optimized Approach to Align Language ModelsXiao Yu, Qingyang Wu, Yu Li, Zhou YuEMNLP 2024
- Flora: Effortless Context Construction to Arbitrary Length and ScaleTianxiang Chen, Zhentao Tan, Xiaofan Bo, Yue Wu 等AAAI 2026 · 被引用 2 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference OptimizationGuanzheng Chen, Xin Li, Michael Shieh, Lidong BingICLR 2025
