ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang
摘要
Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks. Project page: https://jasper0314-huang.github.io/thinkact-vla/ Our main contributions are summarized as follows:
• We propose ThinkAct, a dual-system framework that mutually enhances action execution and visual-grounded embodied reasoning connected by visual latent planning.
• We leverage the visual feedback of goal completion and trajectory alignment as actionaligned rewards to allow long-horizon reasoning grounded in the embodied scene.
• We advance visual latent planning to steer downstream action execution by providing reasoning-enhanced trajectory guidance across diverse environments.
• We demonstrate that our learned reasoning VLA enables capabilities of few-shot adaptation, long-horizon planning, and self-correction across diverse embodied manipulation tasks.
2 Related Works
Recent efforts [19,50,51,30,11] have adapted large language models (LLMs) and vision-language models (VLMs) for action-centric tasks by prompting or post-training on curated instruction-following
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelYihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui 等AAAI 2026 · 被引用 76 次
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic ManipulationYifu Yuan, Haiqin Cui, Yaoting Huang, Yibin Chen 等ICLR 2026 · 被引用 48 次
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong 等CVPR 2026 · 被引用 43 次
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action ModelsJianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu 等ICLR 2026 · 被引用 36 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
相关 Paper
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent PlanningChi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen 等CVPR 2026 · 被引用 24 次
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong 等CVPR 2026
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningFanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You 等ICLR 2026 · 被引用 129 次
- Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action ModelsDianqiao Lei, Lianlei ShanICML 2026
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei 等ACM MM 2025 · 被引用 4 次
