Discrete Adjoint Matching
Oswin So, Brian Karrer, Chuchu Fan, Ricky T. Q. Chen, Guan-Horng Liu
摘要
Computation methods for solving entropy-regularized reward optimization -- a class of problems widely used for fine-tuning generative models -- have advanced rapidly. Among those, Adjoint Matching (AM, Domingo-Enrich et al., 2025) has proven highly effective in continuous state spaces with differentiable rewards. Transferring these practical successes to discrete generative modeling, however, remains particularly challenging and largely unexplored, mainly due to the drastic shift in generative model classes to discrete state spaces, which are nowhere differentiable. In this work, we propose Discrete Adjoint Matching (DAM) -- a discrete variant of AM for fine-tuning discrete generative models characterized by Continuous-Time Markov Chains, such as diffusion-based large language models. The core of DAM is the introduction of discrete adjoint-an estimator of the optimal solution to the original problem but formulated on discrete domains-from which standard matching frameworks can be applied. This is derived via a purely statistical standpoint, in contrast to the control-theoretic viewpoint in AM, thereby opening up new algorithmic opportunities for general adjoint-based estimators. We showcase DAM's effectiveness on synthetic and mathematical reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Proximal Diffusion Neural SamplerWei Guo, Jaemoo Choi, Yuchen Zhu, Molei Tao 等ICLR 2026 · 被引用 19 次
- Bridge Matching Sampler: Scalable Sampling via Generalized Fixed-Point Diffusion MatchingDenis Blessing, Lorenz Richter, Julius Berner, Egor Malitskiy 等ICML 2026 · 被引用 8 次
- Discrete Adjoint Schrödinger Bridge SamplerWei Guo, Yuchen Zhu, Xiaochen Du, Juno Nam 等ICML 2026 · 被引用 3 次
- Unsupervised Diffusion Solver for Combinatorial Optimization via Combinatorial Adjoint MatchingShengyu Feng, Tarun Suresh, Yiming YangICML 2026 · 被引用 1 次
- Learning-to-Optimize via Deep Unfolded FlowsAugustinos Saravanos, Oswin So, H M Sabbir Ahmad, Chuchu FanICML 2026
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal ControlCarles Domingo-Enrich, Michal Drozdzal, Brian Karrer, Ricky T. Q. ChenICLR 2025 · 被引用 2 次
- Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsOussama Zekri, Nicolas BoulléNeurIPS 2025 · 被引用 42 次
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
- Target Concrete Score Matching: A Holistic Framework for Discrete DiffusionRuixiang Zhang, Shuangfei Zhai, Yizhe Zhang, James Thornton 等ICML 2025
- Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement LearningHanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao 等ICML 2025
