Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
Xueyang Yu, Cheng Shi, Yang Wang, Sibei Yang
摘要
Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given ego-streaming video input, an assistant proactively answers diverse, evolving questions at the opportune moment, while maintaining synchronized perception and reasoning. This task embodies three key properties: (1) Proactive Coherence, (2) Just-in-Time Responsiveness, and (3) Synchronized Efficiency. To evaluate and address these properties, we first introduce ESTP-Bench (Ego Streaming Proactive Benchmark) alongside the ESTP-F1 metric-a novel framework designed for their rigorous assessment. Secondly, we propose a comprehensive technical pipeline to enable models to tackle this challenging task. This pipeline comprises: (1) a data engine, (2) a multi-stage training strategy, and (3) a proactive dynamic compression technique. Our proposed model effectively addresses these critical properties while outperforming multiple baselines across diverse online and offline benchmarks. Project Page:https://zhangyl4.github.io/publications/eyes-wide-open/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Vision Function Layer in Multimodal LLMsCheng Shi, Yizhou Yu, Sibei YangNeurIPS 2025 · 被引用 20 次
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsYulin Zhang, Cheng Shi, Sibei YangCVPR 2026 · 被引用 1 次
- RefAny3D: 3D Asset-Referenced Diffusion Models for Image GenerationHanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du 等ICLR 2026 · 被引用 1 次
- Discovering Compositional Hallucinations in LVLMsSibei Yang, Ge Zheng, Jiajin Tang, Jiaye Qian 等NeurIPS 2025 · 被引用 1 次
- Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video UnderstandingKe Ma, Jiaqi Tang, Bin Guo, Xueting Han 等ACL 2026
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Proactive Assistant Dialogue Generation from Streaming Egocentric VideosYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto 等EMNLP 2025 · 被引用 1 次
- Proact-VL: A Proactive VideoLLM for Real-Time AI CompanionsWeicai Yan, Yuhong Dai, Qi Ran, Haodong Li 等ICML 2026 · 被引用 6 次
- StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated CognitionXin Ding, Hao Wu, Yifan Yang, Shiqi Jiang 等ICCV 2025 · 被引用 4 次
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun 等CVPR 2026 · 被引用 28 次
- StreamReady: Learning What to Answer and When in Long Streaming VideosShehreen Azad, Vibhav Vineet, Yogesh S. RawatCVPR 2026 · 被引用 19 次
