Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, Federico Tombari
2025年份
13顶会引用
摘要
Our dataset: EgoTempo Free-form Video Q&A Temporal Event Ordering Q: What does the person do after draining the excess water? Object Counting Q: What is the sequence of actions the person performs with the mug? Multi-Modal LLMs Temporal Understanding Limitations of previous egocentric VideoQA datasets Single-frame Understanding Commonsense Reasoning Q: What is the main purpose of using aluminum foil? Q: What is the status of the microwave before the user gets something from it? Action Sequence Q: How many oranges does the person pick from the tree?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMsWenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang 等ACL 2026 · 被引用 21 次
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 等ICLR 2026 · 被引用 19 次
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and HearingMingfei Chen, Zijun Cui, Xiulong Liu, Jinlin Xiang 等NeurIPS 2025 · 被引用 18 次
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question AnsweringYanjun Li, Yuqian Fu, Tianwen Qian, Qi'ao Xu 等AAAI 2026 · 被引用 14 次
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual ReasoningKailing Li, Qi'ao Xu, Tianwen Qian, Yuqian Fu 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
相关 Paper
- Ego-Grounding for Personalized Question-Answering in Egocentric VideosJunbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela YaoCVPR 2026 · 被引用 7 次
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware ReasoningSARA GHAZANFARI, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy 等CVPR 2026 · 被引用 35 次
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsZiyao Shangguan, Chuhan Li, Yuxuan Ding, Yanan Zheng 等ICLR 2025
- LLaVAction: evaluating and training multi-modal large language models for action understandingHaozhe Qi, Shaokai Ye, Alexander Mathis, Mackenzie W. MathisICLR 2026 · 被引用 3 次
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?Junbo Niu, Yifei Li, Ziyang Miao, Chunjiang Ge 等CVPR 2025
