Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
Chaochen Gao, Xing Wu, Qi Fu, Songlin Hu
摘要
Recent advancements in large language models (LLMs) have highlighted the importance of extending context lengths for handling complex tasks. While traditional methods for training on long contexts often use filtered long documents, these approaches lead to domain imbalances, limiting model performance. To address this, techniques like random document concatenation (Standard) and similarity-based methods (KNN, ICLM) have been developed. However, they either sacrifice semantic coherence or diversity. To balance both aspects, we introduce Quest, a query-centric data synthesis method aggregating semantically relevant yet diverse documents. Quest uses a generative model to predict potential queries for each document, grouping documents with similar queries and keywords. Extensive experiments demonstrate Quest's superior performance on long-context tasks, achieving remarkable results with context lengths of up to 1M tokens and confirming its scalability across various model sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong 等ACL 2025 · 被引用 12 次
- EntropyLong: Effective Long-Context Training via Predictive UncertaintyJunlong Jia, Ziyang Chen, Xing Wu, Chaochen Gao 等ICLR 2026 · 被引用 6 次
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu 等ICLR 2026 · 被引用 5 次
- Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan 等ACL 2025 · 被引用 4 次
- Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement LearningZhuoen Chen, Dongfang Li, Meishan Zhang, Baotian Hu 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- LoongRL: Reinforcement Learning for Advanced Reasoning over Long ContextsSiyuan Wang, Gaokai Zhang, Li Lyna Zhang, Ning Shang 等ICLR 2026 · 被引用 28 次
- QUEST: Query-Aware Sparsity for Efficient Long-Context LLM InferenceJiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao 等ICML 2024 · 被引用 316 次
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMsJunlong Jia, Xing Wu, Chaochen Gao, Ziyang Chen 等AAAI 2026
- Expand, Highlight, Generate: RL-driven Document Generation for Passage RerankingArian Askari, Mohammad Aliannejadi, Chuan Meng, Evangelos Kanoulas 等EMNLP 2023 · 被引用 7 次
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
