FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering
Akhil Kedia, Mohd Abbas Zaidi, Haejun Lee
摘要
Generative models have recently started to outperform extractive models in Open Domain Question Answering, largely by leveraging their decoder to attend over multiple encoded passages and combining their information. However, generative models tend to be larger than extractive models due to the need for a decoder, run slower during inference due to auto-regressive decoder beam search, and their generated output often suffers from hallucinations. We propose to extend transformer encoders with the ability to fuse information from multiple passages, using global representation to provide cross-sample attention over all tokens across samples. Furthermore, we propose an alternative answer span probability calculation to better aggregate answer scores in the global space of all samples. Using our proposed method, we outperform the current stateof-the-art method by 2.5 Exact Match score on the Natural Question dataset while using only 25% of parameters and 35% of the latency during inference, and 4.4 Exact Match on We-bQuestions dataset. When coupled with synthetic data augmentation, we outperform larger models on the TriviaQA dataset as well. The latency and parameter savings of our method make it particularly attractive for open-domain question answering, as these models are often compute-intensive. Global Tokens Global Encoder Layer Question + Passage 1 Passage Encoder Layer Question + Passage 2 Passage Encoder Layer Span Classifier 𝑃𝑃 𝑠𝑠 (𝑠𝑠 1 ) Span Classifier 𝑃𝑃 𝑠𝑠 (𝑠𝑠 2 ) 𝑃𝑃 𝐴𝐴 ("𝑁𝑁𝑁𝑁𝑁𝑁 𝑌𝑌𝑌𝑌𝑌𝑌𝑌𝑌") Σ Multi-Head Attention (Global Encoder Layer) x N Encoder Layers Q KV Contextual Fused Representation Conditional Span Probabilities Global String Probabilities Concat All Passages Multi-Head Attention (Passage Encoder Layer) Q KV Concat Global & Passage
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge GraphJiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang 等ICLR 2024 · 被引用 247 次
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model ReasoningXingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu 等WWW 2025 · 被引用 86 次
- Dense X Retrieval: What Retrieval Granularity Should We Use?Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu 等EMNLP 2024 · 被引用 52 次
- Chain-of-Skills: A Configurable Model for Open-Domain Question AnsweringKaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu 等ACL 2023 · 被引用 12 次
- REANO: Optimising Retrieval-Augmented Reader Models through Knowledge Graph GenerationJinyuan Fang, Zaiqiao Meng, Craig MacDonaldACL 2024 · 被引用 7 次
它引用的顶会 Paper20
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
相关 Paper
- You Only Need One Model for Open-domain Question AnsweringHaejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape 等EMNLP 2022
- KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question AnsweringDonghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu 等ACL 2022 · 被引用 108 次
- FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence SelectionYufei Huang, Xu Han, Maosong SunACL 2024
- UnitedQA: A Hybrid Approach for Open Domain Question AnsweringHao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He 等ACL 2021
- Generation-Augmented Retrieval for Open-Domain Question AnsweringYuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen 等ACL 2021
