DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making
Ziru Wang, Mengmeng Wang, Jade Dai, Teli Ma, Guo-Jun Qi, Yong Liu, Guang Dai, Jingdong Wang
Abstract
Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extracting key spatiotemporal patterns from video that capture visual saliency and temporal evolution, referred to as dynamic representation. Building on this, we introduce DynaMind, a framework that enhances decisionmaking through dynamic reasoning. Specifically, we design an adaptive FrameScorer to evaluate video frames based on semantic consistency and visual saliency, assigning each frame an importance score. These scores are used to filter redundant video content and synthesize compact dynamic representations. Leveraging these representations, we predict critical future dynamics and apply a dynamic-guided policy to generate coherent and context-aware actions. Extensive results demonstrate that DynaMind significantly outperforms the baselines across several simulation benchmarks and real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bed28a72-efea-499e-b37d-d176f6b0594bCited by top-tier papers2
- Agentic Spatio-Temporal Grounding via Collaborative ReasoningHeng Zhao, Yew-Soon Ong, Joey Tianyi ZhouSIGIR 2026 · 1 citation
- DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot ManipulationAlex Wang, Zhiwei Dong, Qicheng Bai, Chenshi Zhang et al.CVPR 2026
Builds on14
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu et al.ICLR 2022 · 827 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen et al.ICLR 2024 · 309 citations
Related papers
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual ReasoningKailing Li, Qi'ao Xu, Tianwen Qian, Yuqian Fu et al.CVPR 2026 · 12 citations
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu et al.AAAI 2024 · 122 citations
- Grounding Physical Concepts of Objects and Events Through Dynamic Visual ReasoningZhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong et al.ICLR 2021 · 13 citations
- DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual DynamicsSilin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler et al.ICML 2026
- DANLI: Deliberative Agent for Following Natural Language InstructionsYichi Zhang, Jianing Yang, Jiayi Pan, Shane Storks et al.EMNLP 2022 · 4 citations
