StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
Bingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
摘要
Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby restricting input flexibility and increasing the number of training parameters. To address these challenges, we propose StitchFusion, a straightforward yet effective modal fusion framework that integrates large-scale pre-trained models directly as encoders and feature fusers. This approach facilitates comprehensive multi-modal and multi-scale feature fusion, accommodating any visual modal inputs. Specifically, our framework achieves modal integration during encoding by sharing multi-modal visual information. To enhance information exchange across modalities, we introduce a multi-directional Modality Adapter module (MoA) to enable cross-modal information transfer during encoding. By leveraging MoA to propagate multi-scale information across pre-trained encoders during the encoding process, StitchFusion achieves multi-modal visual information integration during encoding. Extensive comparative experiments demonstrate that our model achieves state-of-the-art performance on four multi-modal segmentation datasets with minimal additional parameters. Furthermore, the experimental integration of MoA with existing Feature Fusion Modules (FFMs) highlights their complementary nature. Our anonymous code is https://anonymous.4open.science/r/StitchFusion_V2-E777
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationYuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li 等CVPR 2026 · 被引用 31 次
- Exploring Efficient Open-Vocabulary Segmentation in the Remote SensingBingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao 等AAAI 2026 · 被引用 22 次
- UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware PreservationXingyuan Li, Songcheng Du, Yang Zou, Haoyuan Xu 等CVPR 2026 · 被引用 6 次
- Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-TuningJunhao Xiao, Zhiyu Wu, Hao Lin, Yi Chen 等AAAI 2026 · 被引用 4 次
- Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark DatasetSongcheng Du, Yang Zou, Jiaxin Li, Mingxuan Liu 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis 等ACM MM 2022 · 被引用 167 次
- MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationZhiwei Hao, Zhongyu Xiao, Jianyuan Guo, Li Shen 等NeurIPS 2025 · 被引用 1 次
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen 等NeurIPS 2025 · 被引用 8 次
- Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer FusionYikai Wang, Fuchun Sun, Ming Lu, Anbang YaoACM MM 2020 · 被引用 66 次
- Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationJiaxin Cai, Jingze Su, Qi Li, Wenjie Yang 等CVPR 2025
