Lune

NeurIPS2025Top-tier venue

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang, Fu-En Yang

2025Year
179Citations
24Top-tier citations

Abstract

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end fashion, directly mapping inputs to actions without explicit reasoning, which hinders their ability to plan over multiple steps or adapt to complex task variations. In this paper, we propose ThinkAct, a dual-system framework that bridges high-level reasoning with low-level action execution via reinforced visual latent planning. ThinkAct trains a multimodal LLM to generate embodied reasoning plans guided by reinforcing action-aligned visual rewards based on goal completion and trajectory consistency. These reasoning plans are compressed into a visual plan latent that conditions a downstream action model for robust action execution on target environments. Extensive experiments on embodied reasoning and robot manipulation benchmarks demonstrate that ThinkAct enables few-shot adaptation, long-horizon planning, and self-correction behaviors in complex embodied AI tasks. Project page: https://jasper0314-huang.github.io/thinkact-vla/ Our main contributions are summarized as follows:

• We propose ThinkAct, a dual-system framework that mutually enhances action execution and visual-grounded embodied reasoning connected by visual latent planning.

• We leverage the visual feedback of goal completion and trajectory alignment as actionaligned rewards to allow long-horizon reasoning grounded in the embodied scene.

• We advance visual latent planning to steer downstream action execution by providing reasoning-enhanced trajectory guidance across diverse environments.

• We demonstrate that our learned reasoning VLA enables capabilities of few-shot adaptation, long-horizon planning, and self-correction across diverse embodied manipulation tasks.

2 Related Works

Recent efforts [19,50,51,30,11] have adapted large language models (LLMs) and vision-language models (VLMs) for action-centric tasks by prompting or post-training on curated instruction-following

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a4f5248b-d04a-44e4-85e1-e6949db476d0

Cited by top-tier papers24

Ask how each one uses it

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines