Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
Weicai Yan, Wang Lin, Zirun Guo, Ye Wang, Fangming Feng, Xiaoda Yang, Zehan Wang, Tao Jin
Abstract
Prompt learning has demonstrated promising results in fine-tuning pre-trained multimodal models. However, the performance improvement is limited when applied to more complex and fine-grained tasks. The reason is that most existing methods directly optimize the parameters involved in the prompt generation process through loss backpropagation, which constrains the richness and specificity of the prompt representations. In this paper, we propose Diffusion-Driven Prompt Generator (Diff-Prompt), aiming to use the diffusion model to generate rich and fine-grained prompt information for complex downstream tasks. Specifically, our approach consists of three stages. In the first stage, we train a Mask-VAE to compress the masks into latent space. In the second stage, we leverage an improved Diffusion Transformer (DiT) to train a prompt generator in the latent space, using the masks for supervision. In the third stage, we align the denoising process of the prompt generator with the pre-trained model in the semantic space, and use the generated prompts to fine-tune the model. We conduct experiments on a complex pixel-level downstream task, referring expression comprehension, and compare our method with various parameter-efficient fine-tuning approaches. Diff-Prompt achieves a maximum improvement of 8.87 in R@1 and 14.05 in R@5 compared to the foundation model and also outperforms other state-of-the-art methods across multiple metrics. The experimental results validate the effectiveness of our approach and highlight the potential of using generative models for prompt generation. Code is available at https://github.com/Kelvin-ywc/diff-prompt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li et al.ACL 2025 · 16 citations
- Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning ReconstructionTao Wu, Jingyuan Chen, Wang Lin, Jian Zhan et al.ACL 2026 · 3 citations
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkWang Lin, Feng Wang, Majun Zhang, Wentao Hu et al.ICLR 2026 · 2 citations
- Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across DomainsFangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu et al.AAAI 2026 · 1 citation
- Show and Polish: Reference-Guided Identity Preservation in Face Video RestorationWenkang Han, Wang Lin, Yiyun Zhou, Qi Liu et al.ACM MM 2025 · 1 citation
Builds on49
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon et al.NeurIPS 2025 · 6 citations
- In-Context Learning Unlocked for Diffusion ModelsZhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen et al.NeurIPS 2023 · 128 citations
- Masked Diffusion Transformer is a Strong Image SynthesizerShanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng YanICCV 2023 · 290 citations
- Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image GenerationYi Wu, Shengju Qian, Lingting Zhu, Lei Liu et al.CVPR 2026 · 8 citations
- Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language ModelsLiangchen Liu, Nannan Wang, Xi Yang, Xinbo Gao et al.ICML 2025
