MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
Vanya Cohen, Ray Mooney
摘要
Entity state tracking is a necessary component of world modeling that requires maintaining coherent representations of entities over time. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce MET-Bench, a multimodal entity tracking benchmark designed to evaluate the ability of vision-language models to track entity states across modalities. Using three domains, we assess how effectively current models integrate textual and image-based state updates. Our findings reveal a significant performance gap between text-based and image-based entity tracking. We empirically show this discrepancy primarily stems from deficits in visual reasoning rather than perception. We further show that explicit text-based reasoning strategies improve performance, yet limitations remain, especially in long-horizon multimodal tasks. We apply reinforcement learning to improve entity tracking in open-source VLMs. This yields substantial in-modality gains, but does not transfer robustly across input modalities. Our results highlight the need for improved multimodal representations and reasoning techniques to bridge the gap between textual and visual entity tracking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Chess as a Testbed for Language Model State TrackingShubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin GimpelAAAI 2022 · 被引用 77 次
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas 等ICLR 2023 · 被引用 60 次
- A Dataset for Tracking Entities in Open Domain Procedural TextNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal 等EMNLP 2020 · 被引用 38 次
- Towards Coherent and Consistent Use of Entities in Narrative GenerationPinelopi Papalampidi, Kris Cao, Tomás KociskýICML 2022 · 被引用 17 次
相关 Paper
- ProgressLM: Towards Progress Reasoning in Vision-Language ModelsJianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu 等ACL 2026 · 被引用 7 次
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent EnvironmentsZelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan 等CVPR 2026 · 被引用 3 次
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng 等NeurIPS 2025 · 被引用 61 次
