Lune

KDD2026顶会

Towards Efficient Embodied Reasoning: Mixture-of-Depth Compute Allocation for Vision-Language-Action Model

Weiying Xie, Qingchen Zeng, Zihan Meng, Jiayun Tian, Sibo He, Danian Yang, Jie Du, Yunke Wang, Daixun Li, Hengyi Wang, Jitao Ma, Leyuan Fang, Yunsong Li

2026年份

摘要

Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lightweight techniques to compress the backbone. However, these information-lossy methods degrade spatial representations for action generation. In contrast, rate-distortion principles aim to reduce computation while retaining control-sufficient information. Inspired by this insight, we introduce Effective-Edge Flow, an action-aligned attribution measure that quantifies the marginal contribution of token interactions across network depth. This analysis reveals a consistent depth asymmetry, with visual evidence dominating early layers and linguistic reasoning sustaining task-relevant influence into deeper layers. Building on this structure, we propose MoDeVLA, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning. Extensive real-robot evaluations across 20 tasks and multiple embodiments demonstrate that MoDeVLA preserves task performance while reducing latency by about 38% and FLOPs by 86% on edge device NVIDIA Jetson Orin, highlighting its strong ability for embodied systems deployment.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖