VFScale: Intrinsic Reasoning through Verifier-Free Test-time Scalable Diffusion Model
Tao Zhang, Jia-Shu Pan, Ruiqi Feng, Tailin Wu
Abstract
Inspired by human SYSTEM 2 thinking, LLMs excel at complex reasoning tasks via extended Chain-of-Thought. However, similar test-time scaling for diffusion models to tackle complex reasoning remains largely unexplored. From existing work, two primary challenges emerge in this setting: (i) the dependence on an external verifier indicating a notable gap from intrinsic reasoning of human intelligence without any external feedback, and (ii) the lack of an efficient search algorithm. In this paper, we introduce the Verifier-free Test-time Scalable Diffusion Model (VFScale) to achieve scalable intrinsic reasoning, which equips number-of-sample test-time scaling with the intrinsic energy function of diffusion models as the verifier. Concretely, VFScale comprises two key innovations to address the aforementioned challenges. On the training side, VFScale consists of a novel MRNCL loss and a KL regularization to improve the energy landscape, ensuring that the learned energy function itself serves as a reliable verifier. On the inference side, VFScale integrates the denoising process with a novel hybrid Monte Carlo Tree Search (hMCTS) to improve search efficiency. On challenging reasoning tasks of Maze and Sudoku, we demonstrate the effectiveness of VFScale's training objective and scalable inference method. In particular, trained with Maze sizes of up to , our VFScale solves 88% of Maze problems with much larger sizes of , while standard diffusion models completely fail. The code can be found at https://github.com/AI4Science-WestlakeU/VFScale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aefe2502-97ad-48ca-8b65-407fb21659e6Cited by top-tier papers3
- Adaptive Inference-Time Scaling via Cyclic Diffusion SearchGyubin Lee, Bao Truong, Jaesik Yoon, Dongwoo Lee et al.NeurIPS 2025 · 9 citations
- Compositional Monte Carlo Tree Diffusion for Extendable PlanningJaesik Yoon, Hyeonseo Cho, Sungjin AhnNeurIPS 2025 · 2 citations
- Fast Monte Carlo Tree Diffusion: 100× Speedup via Parallel and Sparse PlanningJaesik Yoon, Hyeonseo Cho, Yoshua Bengio, Sungjin AhnNeurIPS 2025
Builds on15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen et al.ACL 2026 · 3 citations
- UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action BranchingKou Misaki, Takuya AkibaICML 2026 · 1 citation
- Test-Time Scaling in Diffusion LLMS via Hidden Semi-Autoregressive ExpertsJihoon Lee, Hoyeon Moon, Kevin Zhai, Arun Kumar Chithanar et al.ICLR 2026 · 7 citations
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward MechanismCongzhi Zhang, Jiawei Peng, Zhenglin Wang, Yilong Lai et al.ACL 2025 · 6 citations
- DiffuReason: Enhancing Reasoning Ability for Diffusion Language Models via Monte Carlo Tree SearchYIPING SONG, Jinyu You, Zhiliang Tian, Jinsong Su et al.ICML 2026
