Lune

ICML2026顶会

Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models

Jiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen, Furong Huang

2026年份
1被引次数

摘要

Masked diffusion language models decode by iteratively unmasking tokens, where the unmasking order defines an ``order of thought'' that strongly influences generation quality yet is typically chosen heuristically. We derive a tractable upper bound on the sequential decoding mismatch, measured by the Kullback–Leibler divergence and expressed in terms of the model’s pathwise log-likelihood, with tightness under sufficient model expressivity. This bound induces a dense self-aware reward for a target sequence xx and unmasking order σ\sigma, over ordered paths, casting order selection as a principled policy optimization problem with a frozen denoiser. We instantiate this idea as Self-Aware Scheduling (SAS), which learns a lightweight order policy using Group Relative Policy Optimization and applies seamlessly to both sequential and semi-autoregressive decoding. On Sudoku with 1B MDM, SAS improves puzzle accuracy from 82.0%82.0\% (best heuristic schedule) to 91.8%91.8\%, and reaches 97.9%97.9\% with second-stage fine-tuning along learned trajectories. On LLaDA-8B, SAS improves pass@1 on GSM8K from 64%64\% to 76%76\% (full diffusion) and on MBPP from 39.5%39.5\% to 41%41\%, while consistently matching or exceeding heuristic schedules across generation lengths and block sizes.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖