Plug-and-Play Context Feature Reuse for Efficient Masked Generation
Xuejie Liu, Anji Liu, Guy Van den Broeck, Yitao Liang
Abstract
Masked generative models (MGMs) have emerged as a powerful framework for image synthesis, combining parallel decoding with strong bidirectional context modeling. However, generating high-quality samples typically requires many iterative decoding steps, resulting in high inference costs. A straightforward way to speed up generation is by decoding more tokens in each step, thereby reducing the total number of steps. However, when many tokens are decoded simultaneously, the model can only estimate the univariate marginal distributions independently, failing to capture the dependency among them. As a result, reducing the number of steps significantly compromises generation fidelity. In this work, we introduce ReCAP (Reused Context-Aware Prediction), a plug-and-play module that accelerates inference in MGMs by constructing low-cost steps via reusing feature embeddings from previously decoded context tokens. ReCAP interleaves standard full evaluations with lightweight steps that cache and reuse context features, substantially reducing computation while preserving the benefits of fine-grained, iterative generation. We demonstrate its effectiveness on top of three representative MGMs (MaskGIT [5], MAGE [27], and MAR [29]), including both discrete and continuous token spaces and covering diverse architectural designs. In particular, on ImageNet256 [7] class-conditional generation, ReCAP achieves up to 2.4× faster inference than the base model with minimal performance drop, and consistently delivers better efficiency-fidelity trade-offs under various generation settings. Our code is publicly available at https://github.com/liebenxj/ReCAP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Lookahead Path Likelihood Optimization for Diffusion LLMsXuejie Liu, Vit Chun Yap, Yitao Liang, Anji LiuICML 2026 · 1 citation
- LazyVAR: Accelerating Visual Autoregressive Models via Scale-wise Token Pruning and Parallel Group DecodingRongge Mao, Chengqi Dong, S Kevin ZhouCVPR 2026
Builds on32
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- Partition Generative Modeling: Masked Modeling Without MasksJustin Deschenaux, Lan Tran, Caglar GulcehreICLR 2026 · 6 citations
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu et al.CVPR 2022 · 346 citations
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage SamplingFeihong Yan, Yao Zhu, Peiru Wang, Pang Kaiyu et al.ICLR 2026 · 3 citations
- LazyMAR: Accelerating Masked Autoregressive Models Via Feature CachingFeihong Yan, Qingyan Wei, Jiayi Tang, Jiajun Li et al.ICCV 2025 · 2 citations
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language ModelsShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.CVPR 2026 · 6 citations
