ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang
Abstract
Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks. Project page: https://jasper0314-huang.github.io/thinkact-vla/ Our main contributions are summarized as follows:
• We propose ThinkAct, a dual-system framework that mutually enhances action execution and visual-grounded embodied reasoning connected by visual latent planning.
• We leverage the visual feedback of goal completion and trajectory alignment as actionaligned rewards to allow long-horizon reasoning grounded in the embodied scene.
• We advance visual latent planning to steer downstream action execution by providing reasoning-enhanced trajectory guidance across diverse environments.
• We demonstrate that our learned reasoning VLA enables capabilities of few-shot adaptation, long-horizon planning, and self-correction across diverse embodied manipulation tasks.
2 Related Works
Recent efforts [19,50,51,30,11] have adapted large language models (LLMs) and vision-language models (VLMs) for action-centric tasks by prompting or post-training on curated instruction-following
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4f5248b-d04a-44e4-85e1-e6949db476d0Cited by top-tier papers24
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.ICLR 2026 · 335 citations
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelYihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui et al.AAAI 2026 · 76 citations
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic ManipulationYifu Yuan, Haiqin Cui, Yaoting Huang, Yibin Chen et al.ICLR 2026 · 48 citations
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong et al.CVPR 2026 · 43 citations
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action ModelsJianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu et al.ICLR 2026 · 36 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
Related papers
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent PlanningChi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen et al.CVPR 2026 · 24 citations
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong et al.CVPR 2026
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningFanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You et al.ICLR 2026 · 129 citations
- Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action ModelsDianqiao Lei, Lianlei ShanICML 2026
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei et al.ACM MM 2025 · 4 citations
