Lune

ICML2025顶会

LOGO - Long cOntext aliGnment via efficient preference Optimization

Zecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu, Min Zhang

出版方
2025年份
8顶会引用

摘要

Long-context models (LCMs) have shown great potential in processing long sequences, with research showing they can accurately locate tokenlevel salient information. Yet, the generation performance of these LCMs is far from satisfactory and might result in misaligned responses, such as hallucinations. To enhance the generation capability, existing works have investigated the effects of data size and quality for both pre-training and instruction tuning stages. Though achieving meaningful improvement, previous methods fall short in either effectiveness or efficiency. In this paper, we introduce LOGO, an efficient and effective training strategy that first introduces preference optimization for long-context alignment. LOGO consists of a reference-free preference optimization strategy and a corresponding efficient data synthesis process. By training with only 0.3B data on a single 8×A800 GPU machine for 16 hours, LOGO allows the Llama-3-8B-Instruct-80K model to achieve comparable performance with GPT-4 in real-world long-context tasks while preserving the model's original capabilities on other tasks, e.g., language modeling and MMLU. Besides, LOGO can also scale the models' context window size while enhancing their performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖