VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip
Wenqi Guo, Shan Du
Abstract
We introduce Value Sign Flip (VSF), a simple and efficient method for incorporating negative prompt guidance in few-step (1-8 steps) diffusion and flow-matching image and video generation models. Unlike existing approaches such as classifier-free guidance (CFG), NASA, and NAG, VSF dynamically suppresses undesired content by flipping the sign of attention values from negative prompts. Our method requires only a small computational overhead and integrates effectively with MMDiT-style architectures such as Stable Diffusion 3.5 Turbo and Flux Schnell, as well as cross-attention-based models like Wan. We validate VSF on a proposed challenging dataset, NegGenBench, with complex prompt pairs. Experimental results on our proposed dataset show that VSF significantly improves negative prompt adherence (reaching 0.420 negative score for quality settings and 0.545 for strong settings) compared to prior methods in few-step models (scored 0.320-0.380 negative score) and even CFG in non-few-step models (scored 0.300 negative score), while maintaining competitive image quality and positive prompt adherence. Our method also suppressed a generate-then-edit pipeline, while also having a much faster runtime. Code, ComfyUI node, and dataset are available in https://github.com/weathon/VSF/tree/main.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39cb26f0-0889-47a9-9132-0cbd58e7542dBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Normalized Attention Guidance: Universal Negative Guidance for Diffusion ModelsDar-Yen Chen, Hmrishav Bandyopadhyay, Kai Zou, Yi-Zhe SongNeurIPS 2025 · 17 citations
- Supercharged One-Step Text-to-Image Diffusion Models with Negative PromptsViet Nguyen, Anh Nguyen, Trung Dao, Khoi Nguyen et al.ICCV 2025 · 1 citation
- Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion ModelsFu-Yun Wang, Yunhao Shui, Jingtan Piao, Keqiang Sun et al.ICLR 2025
- Guiding Diffusion Models with Semantically Degraded ConditionsShilong Han, Yuming Zhang, Hongxia WangCVPR 2026 · 1 citation
- Guiding Diffusion Models With Adaptive Negative Sampling Without External ResourcesAlakh Desai, Nuno VasconcelosICCV 2025 · 1 citation
