SatSynth: Augmenting Image-Mask Pairs Through Diffusion Models for Aerial Semantic Segmentation
Aysim Toker, Marvin Eisenberger, Daniel Cremers, Laura Leal-Taixé
摘要
In recent years, semantic segmentation has become a pivotal tool in processing and interpreting satellite imagery. Yet, a prevalent limitation of supervised learning techniques remains the need for extensive manual annotations by experts. In this work, we explore the potential of generative image diffusion to address the scarcity of annotated data in earth observation tasks. The main idea is to learn the joint data manifold of images and labels, leveraging recent ad-vancements in denoising diffusion probabilistic models. To the best of our knowledge, we are the first to generate both images and corresponding masks for satellite segmentation. We find that the obtained pairs not only display high quality in fine-scale features but also ensure a wide sampling diversity. Both aspects are crucial for earth observation data, where semantic classes can vary severely in scale and occurrence frequency. We employ the novel data instances for downstream segmentation, as a form of data augmentation. In our experiments, we provide comparisons to prior works based on discriminative diffusion models or GANs. We demonstrate that integrating generated samples yields significant quantitative improvements for satellite semantic segmentation - both compared to baselines and when training only on the original data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging ScenesHaonan Wang, Hanyu Zhou, Haoyue Liu, Luxin YanNeurIPS 2025 · 被引用 4 次
- Harnessing the Power of Foundation Models for Accurate Material ClassificationQINGRAN LIN, Fengwei Yang, Chaolun ZhuCVPR 2026 · 被引用 3 次
- Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language TrackingJiawei Ge, Xinyu Zhang, Jiuxin Cao, Xuelin Zhu 等ACM MM 2025 · 被引用 3 次
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic SegmentationYunkai Yang, Yudong Zhang, Kunquan Zhang, Jinxiao Zhang 等CVPR 2026 · 被引用 2 次
- SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesKaiyu Li, Ruixun Liu, Xiangyong Cao, Xueru Bai 等CVPR 2025
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Label-Efficient Semantic Segmentation with Diffusion ModelsDmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov 等ICLR 2022 · 被引用 700 次
- JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation PromotionHaoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang 等AAAI 2026
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 被引用 154 次
- RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change DetectionZian Cao, Wei Wei, Qingshan Gao, Yuanyuan FuCVPR 2026
- Factorized Diffusion Architectures for Unsupervised Image Generation and SegmentationXin Yuan, Michael MaireNeurIPS 2024 · 被引用 4 次
