Reflex: Real-Time Vision-Language-Action Control through Streaming Inference
Yuanchun Guo, Bingyan Liu
摘要
Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow re-computation or mathematically incorrect cache reuse. We present Reflex, a framework that enables real-time streaming inference for flow matching policies by exploiting the Timestep-Invariance Property---that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce AdaRMSNorm, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an async pipeline that decouples visual encoding from action generation, combined with operator fusion that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58 inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% and enabling efficient deployment without performance degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker, Sergey LevineNeurIPS 2025 · 被引用 280 次
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu 等NeurIPS 2025 · 被引用 95 次
- FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency ConsistencyYifei Su, Ning Liu, Dong Chen, Zhen Zhao 等NeurIPS 2025 · 被引用 20 次
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji 等ICLR 2026 · 被引用 60 次
- Block-wise Adaptive Caching for Accelerating Diffusion PolicyKangye Ji, Yuan Meng, Hanyun Cui, Ye Li 等ICLR 2026 · 被引用 9 次
