Sliced Wasserstein Bridge for Open-Vocabulary Video Instance Segmentation
Zheyun Qin, Deng Yu, Chuanchen Luo, Zhumin Chen
Abstract
In recent years, researchers have explored the task of openvocabulary video instance segmentation, which aims to identify, track, and segment any instance within an open set of categories. The core challenge of Open-Vocabulary VIS lies in solving the cross-domain alignment problem, including spatial-temporal and text-visual domain alignments. Existing methods have made progress but still face shortcomings in addressing these alignments, especially due to data heterogeneity. Inspired by metric learning, we propose an innovative Sliced Wasserstein Bridging Learning Framework. This framework utilizes the Sliced Wasserstein distance as the core tool for metric learning, effectively bridging the four domains involved in the task. Our innovations are threefold: (1) Domain Alignment: By mapping features from different domains into a unified metric space, our method maintains temporal consistency and learns intrinsic consistent features between modalities, improving the fusion of text and visual information. (2) Weighting Mechanism: We introduce an importance weighting mechanism to enhance the discriminative ability of our method when dealing with imbalanced or significantly different data. (3) High Efficiency: Our method inherits the computational efficiency of the Sliced Wasserstein distance, allowing for online processing of large-scale video data while maintaining segmentation accuracy. Through extensive experimental evaluations, we have validated the robustness of our concept and the effectiveness of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f0330e7-39ba-4f51-b1c4-e3b4200405a3Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- DVIS: Decoupled Video Instance Segmentation FrameworkTao Zhang, Xingye Tian, Yu Wu, Shunping Ji et al.ICCV 2023 · 86 citations
Related papers
- Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationZesen Cheng, Kehan Li, Hao Li, Peng Jin et al.AAAI 2025 · 3 citations
- Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and UniformityChen-Yi Lu, Kasra Derakhshandeh, Somali ChaterjiCVPR 2025
- OVIS: Open-Vocabulary Visual Instance Search via Visual-Semantic Aligned Representation LearningSheng Liu, Kevin Lin, Lijuan Wang, Junsong Yuan et al.AAAI 2022 · 3 citations
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
- Video Instance Segmentation by Weighted Structure InferenceZheyun Qin, Deng Yu, Yang Shi, Qiangchang Wang et al.ACM MM 2025 · 5 citations
