FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching
Andranik Sargsyan, Shant Navasardyan
摘要
Accurate image segmentation is essential for modern computer vision applications such as image editing, autonomous driving, and medical image analysis. In recent years, Dichotomous Image Segmentation (DIS) has become a standard task for training and evaluating highly accurate segmentation models. Existing DIS approaches often fail to preserve fine-grained details or fully capture the semantic structure of the foreground. To address these challenges, we present FlowDIS, a novel dichotomous image segmentation method built on the flow matching framework, which learns a time-dependent vector field to transport the image distribution to the corresponding mask distribution, optionally conditioned on a text prompt. Moreover, with our Position-Aware Instance Pairing (PAIP) training strategy, FlowDIS offers strong controllability through text prompts, enabling precise, pixel-level object segmentation. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches both with and without language guidance. Compared with the best prior DIS method, FlowDIS achieves a 5.5% higher measure and 43% lower MAE () on the DIS-TE test set. The code is available at: https://github.com/Picsart-AI-Research/FlowDIS
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- LawDIS: Language-Window-Based Controllable Dichotomous Image SegmentationXinyu Yan, Meijun Sun, Ge-Peng Ji, Fahad Shahbaz Khan 等ICCV 2025 · 被引用 3 次
- Shifting the Breaking Point of Flow Matching for Multi-Instance EditingCarmine Zaccagnino, Fabio Quattrini, Enis Simsar, Marta Gazulla 等ICML 2026 · 被引用 1 次
- FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned SegmentationZekang Zhang, Guangyu Gao, YouyunTang, WU CHENGJING 等ICML 2026
- MaskFactory: Towards High-quality Synthetic Data Generation for Dichotomous Image SegmentationHaotian Qian, Yinda Chen, Shengtao Lou, Fahad Shahbaz Khan 等NeurIPS 2024
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image GenerationPaul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret 等AAAI 2026 · 被引用 1 次
