Multi-Field Adaptive Retrieval
Millicent Li, Tongfei Chen, Benjamin Van Durme, Patrick Xia
Abstract
Document retrieval for tasks such as search and retrieval-augmented generation typically involves datasets that are unstructured: free-form text without explicit internal structure in each document. However, documents can have some structure, containing fields such as an article title, a message body, or an HTML header. To address this gap, we introduce Multi-Field Adaptive Retrieval (MFAR), a flexible framework that accommodates any number and any type of document indices on semi-structured data. Our framework consists of two main steps: (1) the decomposition of an existing document into fields, each indexed independently through dense and lexical methods, and (2) learning a model which adaptively predicts the importance of a field by conditioning on the document query, allowing on-the-fly weighting of the most likely field(s). We find that our approach allows for the optimized use of dense versus lexical representations across field types, significantly improves in document ranking over a number of existing retrievers, and achieves state-of-the-art performance for multi-field semi-structured data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e754741-a82f-4674-9119-2a78f45455d1Cited by top-tier papers2
- Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth RetrievalJoaquín Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori et al.ACL 2026 · 1 citation
- Database-Augmented Query Representation for Information RetrievalSoyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang et al.EMNLP 2025
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive ReasoningShirley Wu, Shiyu Zhao, Qian Huang, Kexin Huang et al.NeurIPS 2024 · 95 citations
Related papers
- MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringHui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai et al.ACM MM 2025
- QuDAR: Query-Wise Dual-Perspective Adaptive RetrievalJoeun Kim, Seunghyouk Yoon, Xuan-Bach Le, Youngeun Nam et al.ACL 2026
- Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention ReallocationWenbin An, Jiahao Nie, Feng Tian, Haonan Lin et al.NeurIPS 2025 · 4 citations
- DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive BenchmarkRuofan Hu, Menghui Zhu, Jieming Zhu, Bo Chen et al.KDD 2026 · 1 citation
- GRAD: Generalizing RAG Adaptation with DecodingYoungwon Lee, Seung-won Hwang, Zhewei Yao, Yuxiong HeACL 2026
