Lune

ICLR2025Top-tier venue

Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Tsung-Han Wu, Giscard Biamby, Jerome Quenum, Ritwik Gupta, Joseph E. Gonzalez, Trevor Darrell, David M. Chan

2025Year
16Top-tier citations

Abstract

Large Multimodal Models (LMMs) have made significant strides in visual questionanswering for single images. Recent advancements like long-context LMMs have allowed them to ingest larger, or even multiple, images. However, the ability to process a large number of visual tokens does not guarantee effective retrieval and reasoning for multi-image question answering (MIQA), especially in real-world applications like photo album searches or satellite imagery analysis. In this work, we first assess the limitations of current benchmarks for long-context LMMs. We address these limitations by introducing a new vision-centric, long-context benchmark, "Visual Haystacks (VHs)". We comprehensively evaluate both open-source and proprietary models on VHs, and demonstrate that these models struggle when reasoning across potentially unrelated images, perform poorly on cross-image reasoning, as well as exhibit biases based on the placement of key information within the context window. Towards a solution, we introduce MIRAGE (Multi-Image Retrieval Augmented Generation), an open-source, lightweight visual-RAG framework that processes up to 10k images on a single 40G A100 GPU-far surpassing the 1k-image limit of contemporary models. MIRAGE demonstrates up to 13% performance improvement over existing open-source LMMs on VHs, sets a new state-of-the-art on the RetVQA multi-image QA benchmark, and achieves competitive performance on single-image QA with state-of-the-art LMMs. Our dataset, model, and code are available at: https://visual-haystacks.github.io.

• We introduce a new benchmark, "Visual Haystacks (VHs)" which explicitly tests MIQA models on their ability to retrieve and integrate visual information.

• We systematically assess existing open and closed-source LMMs on the VHs benchmark, and reveal three key findings: susceptibility to visual distractors, difficulty in multi-image reasoning, and a bias in image positioning.

• We introduce a novel baseline for VHs, MIRAGE (Multi-Image Retrieval Augmented Generation), which is the first open-source visual-RAG framework capable of scaling to over 10k images.

The Needle-In-A-Haystack (NIAH) evaluation Kamradt (2023) has recently become a widely adopted unit test for evaluating LLM/LMM systems' ability to uniformly process long-context inputs. In the vision

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 32b18f07-8fe9-4bd7-9d2a-c876779d1c1e

Cited by top-tier papers16

Ask how each one uses it

Builds on29

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines