Lune

ICLR2025顶会

Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Tsung-Han Wu, Giscard Biamby, Jerome Quenum, Ritwik Gupta, Joseph E. Gonzalez, Trevor Darrell, David M. Chan

出版方
2025年份
16顶会引用

摘要

Large Multimodal Models (LMMs) have made significant strides in visual questionanswering for single images. Recent advancements like long-context LMMs have allowed them to ingest larger, or even multiple, images. However, the ability to process a large number of visual tokens does not guarantee effective retrieval and reasoning for multi-image question answering (MIQA), especially in real-world applications like photo album searches or satellite imagery analysis. In this work, we first assess the limitations of current benchmarks for long-context LMMs. We address these limitations by introducing a new vision-centric, long-context benchmark, "Visual Haystacks (VHs)". We comprehensively evaluate both open-source and proprietary models on VHs, and demonstrate that these models struggle when reasoning across potentially unrelated images, perform poorly on cross-image reasoning, as well as exhibit biases based on the placement of key information within the context window. Towards a solution, we introduce MIRAGE (Multi-Image Retrieval Augmented Generation), an open-source, lightweight visual-RAG framework that processes up to 10k images on a single 40G A100 GPU-far surpassing the 1k-image limit of contemporary models. MIRAGE demonstrates up to 13% performance improvement over existing open-source LMMs on VHs, sets a new state-of-the-art on the RetVQA multi-image QA benchmark, and achieves competitive performance on single-image QA with state-of-the-art LMMs. Our dataset, model, and code are available at: https://visual-haystacks.github.io.

• We introduce a new benchmark, "Visual Haystacks (VHs)" which explicitly tests MIQA models on their ability to retrieve and integrate visual information.

• We systematically assess existing open and closed-source LMMs on the VHs benchmark, and reveal three key findings: susceptibility to visual distractors, difficulty in multi-image reasoning, and a bias in image positioning.

• We introduce a novel baseline for VHs, MIRAGE (Multi-Image Retrieval Augmented Generation), which is the first open-source visual-RAG framework capable of scaling to over 10k images.

The Needle-In-A-Haystack (NIAH) evaluation Kamradt (2023) has recently become a widely adopted unit test for evaluating LLM/LMM systems' ability to uniformly process long-context inputs. In the vision

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper16

问问它们各自怎么用它

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖