Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
Xueyang Yu, Cheng Shi, Yang Wang, Sibei Yang
Abstract
Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given ego-streaming video input, an assistant proactively answers diverse, evolving questions at the opportune moment, while maintaining synchronized perception and reasoning. This task embodies three key properties: (1) Proactive Coherence, (2) Just-in-Time Responsiveness, and (3) Synchronized Efficiency. To evaluate and address these properties, we first introduce ESTP-Bench (Ego Streaming Proactive Benchmark) alongside the ESTP-F1 metric-a novel framework designed for their rigorous assessment. Secondly, we propose a comprehensive technical pipeline to enable models to tackle this challenging task. This pipeline comprises: (1) a data engine, (2) a multi-stage training strategy, and (3) a proactive dynamic compression technique. Our proposed model effectively addresses these critical properties while outperforming multiple baselines across diverse online and offline benchmarks. Project Page:https://zhangyl4.github.io/publications/eyes-wide-open/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d559a70-2e4a-4a5a-9fe0-dcb92c8ff968Cited by top-tier papers5
- Vision Function Layer in Multimodal LLMsCheng Shi, Yizhou Yu, Sibei YangNeurIPS 2025 · 20 citations
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsYulin Zhang, Cheng Shi, Sibei YangCVPR 2026 · 1 citation
- RefAny3D: 3D Asset-Referenced Diffusion Models for Image GenerationHanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du et al.ICLR 2026 · 1 citation
- Discovering Compositional Hallucinations in LVLMsSibei Yang, Ge Zheng, Jiajin Tang, Jiaye Qian et al.NeurIPS 2025 · 1 citation
- Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video UnderstandingKe Ma, Jiaqi Tang, Bin Guo, Xueting Han et al.ACL 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Proactive Assistant Dialogue Generation from Streaming Egocentric VideosYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto et al.EMNLP 2025 · 1 citation
- Proact-VL: A Proactive VideoLLM for Real-Time AI CompanionsWeicai Yan, Yuhong Dai, Qi Ran, Haodong Li et al.ICML 2026 · 6 citations
- StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated CognitionXin Ding, Hao Wu, Yifan Yang, Shiqi Jiang et al.ICCV 2025 · 4 citations
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun et al.CVPR 2026 · 28 citations
- StreamReady: Learning What to Answer and When in Long Streaming VideosShehreen Azad, Vibhav Vineet, Yogesh S. RawatCVPR 2026 · 19 citations
