Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
Arsha Nagrani, Jasper Uijlings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A. Ross, Cordelia Schmid
Abstract
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva-Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high-quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi-step multimodal questions and spatiotemporally-dense human-annotated reasoning traces. Benchmarking experiments show that state-of-the-art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatio-temporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva-Ego can be downloaded at https: //github.com/google-deepmind/neptune
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1bfa4d1-142e-4af2-9fda-b1723a12e608Builds on10
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng et al.ICCV 2025 · 28 citations
- Egocentric Planning for Scalable Embodied Task AchievementXiaotian Liu, Héctor Palacios, Christian MuiseNeurIPS 2023 · 9 citations
- Minerva: Evaluating Complex Video ReasoningArsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch et al.ICCV 2025 · 5 citations
Related papers
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint FramesSahithya Ravi, Gabriel Herbert Sarch, Vibhav Vineet, Andrew D. Wilson et al.EMNLP 2025 · 1 citation
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li et al.ICLR 2026 · 22 citations
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu et al.ICLR 2026 · 53 citations
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in RoboticsSimindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai et al.AAAI 2026 · 2 citations
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View ScenesMohsen Gholami, Ahmad Rezaei, Zhou Weimin, Sitong Mao et al.ICLR 2026 · 67 citations
