From Predictors to Samplers via the Training Trajectory
Soumya Ram, Akhila Ram
Abstract
Sampling from trained predictors is fundamental for interpretability and as a compute-light alternative to diffusion models, but local samplers struggle on the rugged, high-frequency functions such models learn. We observe that standard neural‑network training implicitly produces a coarse‑to‑fine sequence of models. Early checkpoints suppress high‑degree/ high‑frequency components (Boolean monomials; spherical harmonics under NTK), while later checkpoints restore detail. We exploit this by running a simple annealed sampler across the training trajectory, using early checkpoints for high‑mobility proposals and later ones for refinement. In the Boolean domain, this can turn the exponential bottleneck arising from rugged landscapes or needle gadgets into a near-linear one. In the continuous domain, under the NTK regime, this corresponds to smoothing under the NTK kernel. Requiring no additional compute, our method shows strong empirical gains across a variety of synthetic and real-world tasks, including constrained sampling tasks that diffusion models are unable to handle.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ef36e02-cec1-46b0-a439-5bbc8775d6f6Builds on22
- Practical and Asymptotically Exact Conditional Sampling in Diffusion ModelsLuhuan Wu, Brian L. Trippe, Christian A. Naesseth, David M. Blei et al.NeurIPS 2023 · 276 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- Design-Bench: Benchmarks for Data-Driven Offline Model-Based OptimizationBrandon Trabucco, Xinyang Geng, Aviral Kumar, Sergey LevineICML 2022 · 126 citations
Related papers
- Progressive Inference-Time Annealing of Diffusion Models for Sampling from Boltzmann DensitiesTara Akhound-Sadegh, Jungyoon Lee, Joey Bose, Valentin De Bortoli et al.NeurIPS 2025 · 28 citations
- Diffusing Differentiable RepresentationsYash Savani, Marc Finzi, J. Zico KolterNeurIPS 2024 · 1 citation
- Accelerated Parallel Tempering via Neural TransportsLeo Zhang, Peter Potaptchik, Jiajun He, Yuanqi Du et al.ICLR 2026 · 14 citations
- DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant ConvolutionZhizhen Wu, Zhe Cao, Yuchi HuoICLR 2026
- NTK-Guided Implicit Neural TeachingChen Zhang, Wei Zuo, Bingyang Cheng, Yikun Wang et al.CVPR 2026 · 3 citations
