Plug-and-Play Context Feature Reuse for Efficient Masked Generation
Xuejie Liu, Anji Liu, Guy Van den Broeck, Yitao Liang
摘要
Masked generative models (MGMs) have emerged as a powerful framework for image synthesis, combining parallel decoding with strong bidirectional context modeling. However, generating high-quality samples typically requires many iterative decoding steps, resulting in high inference costs. A straightforward way to speed up generation is by decoding more tokens in each step, thereby reducing the total number of steps. However, when many tokens are decoded simultaneously, the model can only estimate the univariate marginal distributions independently, failing to capture the dependency among them. As a result, reducing the number of steps significantly compromises generation fidelity. In this work, we introduce ReCAP (Reused Context-Aware Prediction), a plug-and-play module that accelerates inference in MGMs by constructing low-cost steps via reusing feature embeddings from previously decoded context tokens. ReCAP interleaves standard full evaluations with lightweight steps that cache and reuse context features, substantially reducing computation while preserving the benefits of fine-grained, iterative generation. We demonstrate its effectiveness on top of three representative MGMs (MaskGIT [5], MAGE [27], and MAR [29]), including both discrete and continuous token spaces and covering diverse architectural designs. In particular, on ImageNet256 [7] class-conditional generation, ReCAP achieves up to 2.4× faster inference than the base model with minimal performance drop, and consistently delivers better efficiency-fidelity trade-offs under various generation settings. Our code is publicly available at https://github.com/liebenxj/ReCAP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Lookahead Path Likelihood Optimization for Diffusion LLMsXuejie Liu, Vit Chun Yap, Yitao Liang, Anji LiuICML 2026 · 被引用 1 次
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
它引用的顶会 Paper32
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- Partition Generative Modeling: Masked Modeling Without MasksJustin Deschenaux, Lan Tran, Caglar GulcehreICLR 2026 · 被引用 6 次
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu 等CVPR 2022 · 被引用 346 次
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage SamplingFeihong Yan, Yao Zhu, Peiru Wang, Pang Kaiyu 等ICLR 2026 · 被引用 3 次
- LazyMAR: Accelerating Masked Autoregressive Models Via Feature CachingFeihong Yan, Qingyan Wei, Jiayi Tang, Jiajun Li 等ICCV 2025 · 被引用 2 次
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language ModelsShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 等CVPR 2026 · 被引用 6 次
