ICML2026
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
Boyang Xu, Qing Zou, Siqin Yang, Hao Yan
摘要
Distributional reinforcement learning (DRL) models the full return distribution, but typically relies on finite-dimensional categorical or quantile approximations, often involving projection or quantile-regression approximations to the Bellman target, together with independently sampled bootstrap targets that obscure transport structure and add variance. We present Path-Coupled Bellman Flows (PCBF), a continuous-time DRL method that encodes Bellman endpoint consistency and pathwise Bellman-coupled geometry within generative flow trajectories. PCBF represents return distributions via flow matching and couples the paths of consecutive states through shared base noise, yielding a geometric Bellman flow relation between velocity fields. This structure enables a -parameterized control-variate target that reduces training variance while preserving the source and Bellman endpoint geometry. Experiments on analytically tractable MRPs, OGBench, and D4RL show improved distributional fidelity, training stability, and competitive offline RL performance.