Large Language Models to Diffusion Finetuning
Edoardo Cetin, Tianyu Zhao, Yujin Tang
Abstract
We propose a new finetuning method to provide pre-trained large language models (LMs) the ability to scale test-time compute through the diffusion framework. By increasing the number of diffusion steps, we show our finetuned models achieve monotonically increasing accuracy, directly translating to improved performance across downstream tasks. Furthermore, our finetuned models can expertly answer questions on specific topics by integrating powerful guidance techniques, and autonomously determine the compute required for a given problem by leveraging adaptive ODE solvers. Our method is applicable to any foundation model pre-trained with cross-entropy and does not modify any of its original weights, fully preserving its strong single-step generation capabilities. We show our method can be more effective and is fully compatible with traditional finetuning and search approaches, introducing an orthogonal new direction to unify the strengths of the autoregressive and diffusion frameworks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75b9187a-2fc4-43fc-b24d-d74074010b5aCited by top-tier papers5
- CANDI: Hybrid Discrete-Continuous Diffusion ModelsPatrick Pynadath, Jiaxin Shi, Ruqi ZhangICML 2026 · 28 citations
- LaDiR: Latent Diffusion Enhances LLMs for Text ReasoningHaoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Nicklas Majamaki et al.ICLR 2026 · 25 citations
- Watermarking Diffusion Language ModelsThibaud Gloaguen, Robin Staab, Nikola Jovanović, Martin VechevICLR 2026 · 13 citations
- Reinforcement Learning Teachers of Test Time ScalingEdoardo Cetin, Tianyu Zhao, Yujin TangNeurIPS 2025 · 12 citations
- Non-Markovian Discrete Diffusion with Causal Language ModelsYangtian Zhang, Sizhuang He, Daniel LeVine, Lawrence Zhao et al.NeurIPS 2025 · 5 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- TESS 2: A Large-Scale Generalist Diffusion Language ModelJaesung Tae, Hamish Ivison, Sachin Kumar, Arman CohanACL 2025 · 19 citations
- UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action BranchingKou Misaki, Takuya AkibaICML 2026 · 1 citation
- Controlling Text-to-Image Diffusion by Orthogonal FinetuningZeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue et al.NeurIPS 2023 · 277 citations
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for ReasoningCharlie Victor Snell, Jaehoon Lee, Kelvin Xu, Aviral KumarICLR 2025
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 191 citations
