OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
Ding Zhong, Xu Zheng, Chenfei Liao, Yuanhuiyi Lyu, Jialei Chen, Shengyang Wu, Linfeng Zhang, Xuming Hu
摘要
Segment Anything Model 2 (SAM2) has emerged as a strong base model in various pinhole imaging segmentation tasks. However, when applying it to 360° domain, the significant field-of-view (FoV) gap between pinhole and panoramic images poses unique challenges. Two major concerns for this application includes 1) inevitable distortion and object deformation brought by the large FoV disparity between domains; 2) the lack of pixellevel semantic understanding that the original SAM2 cannot provide. To address these issues, we propose a novel OmniSAM framework, which makes the first attempt to apply SAM2 for panoramic semantic segmentation. Specifically, to bridge the first gap, OmniSAM first divides the panorama into sequences of patches. These patches are then treated as image sequences in similar manners as in video segmentation tasks. We then leverage the SAM2's memory mechanism to extract cross-patch correspondences that embeds the cross-FoV dependencies, improving feature continuity and the prediction consistency along mask boundaries. For the second gap, OmniSAM fine-tunes the pretrained image encoder and reutilize the mask decoder for semantic prediction. An FoV-based prototypical adaptation module with dynamic pseudo label update mechanism is also introduced to facilitate the alignment of memory and backbone features, thereby improving model generalization ability across different sizes of source models. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods by large margins, e.g., on SPin8-to-SPan8, on CS13-to-DP13.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement LearningZekai Lin, Xu ZhengCVPR 2026 · 被引用 7 次
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy 等NeurIPS 2025 · 被引用 3 次
- Reducing Unimodal Bias in Multi-Modal Semantic Segmentation With Multi-Scale Functional Entropy RegularizationXu Zheng, Yuanhuiyi Lyu, Lutao Jiang, Danda Pani Paudel 等ICCV 2025 · 被引用 2 次
- Seeing Beyond: Extrapolative Domain Adaptive Panoramic SegmentationYuanfan Zheng, Kunyu Peng, Xu Zheng, Kailun YangCVPR 2026 · 被引用 1 次
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Hiera: A Hierarchical Vision Transformer without the Bells-and-WhistlesChaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei 等ICML 2023 · 被引用 388 次
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationJiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß 等CVPR 2022 · 被引用 100 次
相关 Paper
- Semantics, Distortion, and Style Matter: Towards Source-Free UDA for Panoramic SegmentationXu Zheng, Pengyuan Zhou, Athanasios V. Vasilakos, Lin WangCVPR 2024 · 被引用 16 次
- Unlocking Constraints: Source-Free Occlusion-Aware Seamless SegmentationYihong Cao, Jiaming Zhang, Xu Zheng, Hao Shi 等ICCV 2025 · 被引用 4 次
- OFL-SAM2: Prompt SAM2 with Online Few-shot Learner for Efficient Medical Image SegmentationMeng Lan, Lefei Zhang, Xiaomeng LiAAAI 2026
- Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic SegmentationXu Zheng, Jinjing Zhu, Yexin Liu, Zidong Cao 等CVPR 2023
- GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic SegmentationWeiming Zhang, Yexin Liu, Xu Zheng, Lin WangCVPR 2024 · 被引用 14 次
