EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
Yanjun Li, Yuqian Fu, Tianwen Qian, Qi'ao Xu, Silong Dai, Danda Pani Paudel, Luc Van Gool, Xiaoling Wang
摘要
Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce EgoCross, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. EgoCross covers four diverse and challenging domains, including surgery, industry, extreme sports, and animal perspective, representing realistic and high-impact application scenarios. It comprises approximately 1,000 QA pairs across 798 video clips, spanning four key QA tasks: prediction, recognition, localization, and counting. Each QA pair provides both OpenQA and CloseQA formats to support fine-grained evaluation. Extensive experiments show that most existing MLLMs, whether general-purpose or egocentric-specialized, struggle to generalize to domains beyond daily life, highlighting the limitations of current models. Furthermore, we conduct several pilot studies, e.g., fine-tuning and reinforcement learning, to explore potential improvements. We hope EgoCross and our accompanying analysis will serve as a foundation for advancing domain-adaptive, robust egocentric video understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 等ICLR 2026 · 被引用 19 次
- EgoSound: Benchmarking Sound Understanding in Egocentric VideosBingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 等CVPR 2026 · 被引用 7 次
- PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement LearningZekai Lin, Xu ZhengCVPR 2026 · 被引用 7 次
- Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity EnvironmentsDi Wen, Lei Qi, Kunyu Peng, Kailun Yang 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 等ICCV 2023 · 被引用 152 次
- BlockMix: Meta Regularization and Self-Calibrated Inference for Metric-Based Meta-LearningHao Tang, Zechao Li, Zhimao Peng, Jinhui TangACM MM 2020 · 被引用 120 次
- Adversarial Cross-Domain Action Recognition with Co-AttentionBoxiao Pan, Zhangjie Cao, Ehsan Adeli, Juan Carlos NieblesAAAI 2020 · 被引用 114 次
- Interactive Prototype Learning for Egocentric Action RecognitionXiaohan Wang, Linchao Zhu, Heng Wang, Yi YangICCV 2021 · 被引用 78 次
相关 Paper
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding 等CVPR 2026
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang 等ICML 2026 · 被引用 5 次
- Ego-Grounding for Personalized Question-Answering in Egocentric VideosJunbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela YaoCVPR 2026 · 被引用 7 次
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan 等CVPR 2026 · 被引用 2 次
- EAGLE: Egocentric AGgregated Language-video EngineJing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi 等ACM MM 2024 · 被引用 3 次
