SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
Chaoyang Wang, Xiangtai Li, Lu Qi, Henghui Ding, Yunhai Tong, Ming-Hsuan Yang
Abstract
Semantic segmentation and semantic image synthesis are two representative tasks in visual perception and generation. While existing methods consider them as two distinct tasks, we propose a unified framework (SemFlow) and model them as a pair of reverse problems. Specifically, motivated by rectified flow theory, we train an ordinary differential equation (ODE) model to transport between the distributions of real images and semantic masks. As the training object is symmetric, samples belonging to the two distributions, images and semantic masks, can be effortlessly transferred reversibly. For semantic segmentation, our approach solves the contradiction between the randomness of diffusion outputs and the uniqueness of segmentation results. For image synthesis, we propose a finite perturbation approach to enhance the diversity of generated results without changing the semantic categories. Experiments show that our SemFlow achieves competitive results on semantic segmentation and semantic image synthesis tasks. We hope this simple framework will motivate people to rethink the unification of low-level and high-level vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3dd717b-5f1f-46a1-b047-26ad41dafb7eCited by top-tier papers12
- MotionBooth: Motion-Aware Customized Text-to-Video GenerationJianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang et al.NeurIPS 2024 · 114 citations
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu et al.CVPR 2026 · 26 citations
- FMPose3D: monocular 3D pose estimation via flow matchingTi Wang, Xiaohang Yu, Mackenzie Weygandt MathisCVPR 2026 · 6 citations
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon et al.NeurIPS 2025 · 6 citations
- HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World DehazingJunseong Shin, Seungwoo Chung, Yunjeong Yang, Tae Hyun KimICCV 2025 · 5 citations
Builds on52
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative ModelsFrancisco Caetano, Christiaan G. A. Viviers, Peter H. N. de With, Fons van der SommenAAAI 2026 · 4 citations
- Reconciling Visual Perception and Generation in Diffusion ModelsLiulei Li, Yi Yang, Wenguan WangICLR 2026
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified FlowXingchao Liu, Chengyue Gong, Qiang LiuICLR 2023 · 75 citations
- One Diffusion to Generate Them AllDuong H. Le, Tuan Pham, Sangho Lee, Christopher Clark et al.CVPR 2025
- Scaling Properties of Diffusion Models For Perceptual TasksRahul Ravishankar, Zeeshan Patel, Jathushan Rajasegaran, Jitendra MalikCVPR 2025
