Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering
Avi Caciularu, Matthew E. Peters, Jacob Goldberger, Ido Dagan, Arman Cohan
摘要
The integration of multi-document pre-training objectives into language models has resulted in remarkable improvements in multi-document downstream tasks. In this work, we propose extending this idea by pre-training a generic multi-document model from a novel crossdocument question answering pre-training objective. To that end, given a set (or cluster) of topically-related documents, we systematically generate semantically-oriented questions from a salient sentence in one document and challenge the model, during pre-training, to answer these questions while "peeking" into other topically-related documents. In a similar manner, the model is also challenged to recover the sentence from which the question was generated, again while leveraging cross-document information. This novel multidocument QA formulation directs the model to better recover cross-text informational relations, and introduces a natural augmentation that artificially increases the pre-training data. Further, unlike prior multi-document models that focus on either classification or summarization tasks, our pre-training objective formulation enables the model to perform tasks that involve both short text generation (e.g., QA) and long text generation (e.g., summarization). Following this scheme, we pre-train our model -termed QAMDEN -and evaluate its performance across several multi-document tasks, including multi-document QA, summarization, and query-focused summarization, yielding improvements of up to 7%, and significantly outperforms zero-shot GPT-3.5 and GPT-4. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue 等ICML 2024 · 被引用 204 次
- MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingGabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor 等ACL 2025 · 被引用 13 次
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context InstructionsChaochen Gao, Xing Wu, Zijia Lin, Debing Zhang 等NeurIPS 2025 · 被引用 8 次
- Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language ModelsJunfeng Tian, Da Zheng, Yang Chen, Rui Wang 等ACL 2025 · 被引用 8 次
- SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMsYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 等ACL 2024 · 被引用 6 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
相关 Paper
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 被引用 147 次
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 被引用 463 次
- Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationArtidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng WuACL 2023 · 被引用 4 次
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
- M3: A Multi-View Fusion and Multi-Decoding Network for Multi-Document Reading ComprehensionLiang Wen, Houfeng Wang, Yingwei Luo, Xiaolin WangEMNLP 2022 · 被引用 3 次
