ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Xiaomeng ZHU, Fengming ZHU, Weijie Zhou, Ye Tian, Zhenlin Hu, Yufei Huang, Yuchun Guo, Xinyu Wu, Zhengyou Zhang, Fangzhen Lin, Xuantang Xiong
Abstract
While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To address this, we introduce ProAct-75 , a benchmark designed to train and evaluate proactive agents across diverse domains, including assistance, maintenance, and safety monitoring. Spanning 75 tasks, our dataset features 91,581 step-level annotations enriched with explicit task graphs. These graphs encode step dependencies and parallel execution possibilities, providing the structural grounding necessary for complex decision-making. Building on this benchmark, we propose ProAct-Helper , a reference baseline powered by a Multimodal Large Language Model (MLLM) that grounds decision-making in state detection, and leveraging task graphs to enable entropy-driven heuristic search for action selection, allowing agents to execute parallel threads independently rather than mirroring the human's next step. Extensive experiments demonstrate that ProAct-Helper outperforms strong closed-source models, improving trigger detection mF1 by 6.21%, saving 0.25 more steps in online one-step decision, and increasing the rate of parallel actions by 15.58%. Code is available at https://github.com/ZhuXMMM/ProAct.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 119eb7f5-bbc3-45c8-b325-18fe0856d4d9Builds on14
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo et al.ICCV 2019 · 182 citations
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsJiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu et al.NeurIPS 2023 · 180 citations
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
- ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory PerceptionsBufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu et al.NeurIPS 2025 · 68 citations
Related papers
- Proactive Agent: Shifting LLM Agents from Reactive Responses to Active AssistanceYaxi Lu, Shenzhi Yang, Cheng Qian, Guirong Chen et al.ICLR 2025
- ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile DevicesDezhi Kong, Zhengzhao Feng, Qiliang Liang, Hao Wang et al.CVPR 2026 · 6 citations
- Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI ActionsXiaokai Yan, Jingtao Ding, Yong Li, Zhiwen YuUbiComp 2026
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas et al.ICLR 2025
- TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal ModelsZhenkun Gao, Xuhong Wang, Xin Tan, Yuan XieICLR 2026 · 1 citation
