MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging Videos
Minghan Li, Shuai Li, Wangmeng Xiang, Lei Zhang
摘要
While impressive progress has been achieved, video instance segmentation (VIS) methods with per-clip input often fail on challenging videos with occluded objects and crowded scenes. This is mainly because instance queries in these methods cannot encode well the discriminative embeddings of instances, making the query-based segmenter difficult to distinguish those 'hard' instances. To address these issues, we propose to mine discriminative query embeddings (MDQE) to segment occluded instances on challenging videos. First, we initialize the positional embeddings and content features of object queries by considering their spatial contextual information and the inter-frame object motion. Second, we propose an inter-instance mask repulsion loss to distance each instance from its nearby non-target instances. The proposed MDQE is the first VIS method with per-clip input that achieves state-of-the-art results on challenging videos and competitive performance on simple videos. In specific, MDQE with ResNet50 achieves 33.0% and 44.5% mask AP on OVIS and YouTube-VIS 2021, respectively. Code of MDQE can be found at https: //github.com/MinghanLi/MDQE_CVPR2023 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SeeSR: Towards Semantics-Aware Real-World Image Super-ResolutionRongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang 等CVPR 2024 · 被引用 119 次
- DVIS: Decoupled Video Instance Segmentation FrameworkTao Zhang, Xingye Tian, Yu Wu, Shunping Ji 等ICCV 2023 · 被引用 86 次
- CAVIS: Context-Aware Video Instance SegmentationSeunghun Lee, Jiwan Seo, Kiljoon Han, Minwoo Choi 等ICCV 2025 · 被引用 4 次
- Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance SegmentationZesen Cheng, Kehan Li, Hao Li, Peng Jin 等AAAI 2025 · 被引用 3 次
- LOMM: Latest Object Memory Management for Temporally Consistent Video Instance SegmentationSeunghun Lee, Jiwan Seo, Minwoo Choi, Kiljoon Han 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
相关 Paper
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 被引用 135 次
- A Generalized Framework for Video Instance SegmentationMiran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim 等CVPR 2023
- SyncVIS: Synchronized Video Instance SegmentationRongkun Zheng, Lu Qi, Xi Chen, Yi Wang 等NeurIPS 2024 · 被引用 8 次
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia 等NeurIPS 2022 · 被引用 23 次
- CTVIS: Consistent Training for Online Video Instance SegmentationKaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang 等ICCV 2023 · 被引用 72 次
