Reflex: Real-Time Vision-Language-Action Control through Streaming Inference
Yuanchun Guo, Bingyan Liu
Abstract
Flow matching Vision-Language-Action (VLA) models promise precise continuous control, but their iterative denoising nature introduces fundamental incompatibilities with real-time robotics: global timestep injection invalidates KV-caching, forcing a choice between slow re-computation or mathematically incorrect cache reuse. We present Reflex, a framework that enables real-time streaming inference for flow matching policies by exploiting the Timestep-Invariance Property---that perception encoders are functionally independent of the denoising loop. Reflex partitions the attention context into static, sliding, and dynamic regions, enabling incremental cache updates while preserving full-batch-equivalent attention outputs for fixed inputs. To ensure stability under continuous high-frequency inference, we introduce AdaRMSNorm, an adaptive normalization layer that prevents BFloat16 numerical collapse by gating on flow phase. We further maximize throughput through an async pipeline that decouples visual encoding from action generation, combined with operator fusion that reduces kernel overhead. On LIBERO and Kinetix benchmarks, Reflex achieves a 2.58 inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% and enabling efficient deployment without performance degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9e75b8d-eab2-4a3e-a440-d1e94ce01a1bBuilds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker, Sergey LevineNeurIPS 2025 · 280 citations
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
- FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency ConsistencyYifei Su, Ning Liu, Dong Chen, Zhen Zhao et al.NeurIPS 2025 · 20 citations
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji et al.ICLR 2026 · 60 citations
- Block-wise Adaptive Caching for Accelerating Diffusion PolicyKangye Ji, Yuan Meng, Hanyun Cui, Ye Li et al.ICLR 2026 · 9 citations
