Lune

ICML2026顶会

Meta Flow Maps enable scalable reward alignment

Peter Potaptchik, Adhi Saravanan, Abbas Mammadov, Alvaro Prat, Michael Albergo, Yee-Whye Teh

2026年份
25被引次数
8顶会引用

摘要

Controlling generative models is computationally expensive. This is because optimal alignment with a reward function-whether via inference-time steering or fine-tuning-requires estimating the value function. This task demands access to the conditional posterior p 1|t (x1|xt), the distribution of clean data x1 consistent with an intermediate state xt, a requirement that typically compels methods to resort to costly trajectory simulations. To address this bottleneck, we introduce Meta Flow Maps (MFMs), a framework extending consistency models and flow maps into the stochastic regime. MFMs are trained to perform stochastic one-step posterior sampling, generating arbitrarily many i.i.d. draws of clean data x1 from any intermediate state. Crucially, these samples provide a differentiable reparametrisation that unlocks efficient value function estimation. We leverage this capability to solve bottlenecks in both paradigms: enabling inference-time steering without inner rollouts, and facilitating unbiased, off-policy fine-tuning to general rewards. Empirically, our single-particle steered-MFM sampler outperforms a Best-of-1000 baseline on ImageNet across multiple rewards at a fraction of the compute.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper36

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖