Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels
Junda Xu, Yanmeng Liu, Xiangqiang Zeng, Jinrong Wu, Ying Qu, Libao Zhang
Abstract
Google Earth imagery, combined with building footprint databases, offers an efficient way to construct localized building datasets. However, the lack of orthorectification in these images leads to spatial misalignments between annotations and their corresponding roof locations. Adopting such misaligned data directly for model training can severely degrade segmentation performance. To address the challenge, we propose an Object-based Multi-stage Alignment Framework (OMAF) that generates high-quality corrected labels with minimal manual intervention. OMAF first employs a prior-regularized self-alignment method to produce high-confidence, object-level offset pseudo-labels, which are then used to train an instance-level offset regression model for label refinement. Experimental results on the challenging datasets demonstrate that OMAF effectively corrects misalignments and consistently boosts the mIoU of all baseline models by up to 40.6%. This work provides a practical and cost-effective solution for large-scale dataset construction and domain adaptation. Our code and dataset are publicly available at https://github.com/dayunyan/ OMAF-Building-Alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d079d740-bd83-4067-a654-c974aa688afeBuilds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- Scene Grounding in the WildTamir Cohen, Leo Segre, Shay Shomer Chai, Shai Avidan et al.CVPR 2026 · 1 citation
- Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object DetectionChen Chen, Jiahao Qi, Xingyue Liu, Kangcheng Bin et al.CVPR 2024
- 3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level SupervisionsWeijia Li, Haote Yang, Zhenghao Hu, Juepeng Zheng et al.CVPR 2024
- OmniCity: Omnipotent City Understanding with Multi-Level and Multi-View ImagesWeijia Li, Yawen Lai, Linning Xu, Yuanbo Xiangli et al.CVPR 2023
- Domain-Specific Alignment Network for Multi-Domain Image-Based 3D Object RetrievalYuting Su, Yuqian Li, Dan Song, Zhendong Mao et al.ACM MM 2020 · 3 citations
