EOV-Seg: Efficient Open-Vocabulary Panoptic Segmentation
Hongwei Niu, Jie Hu, Jianghang Lin, Guannan Jiang, Shengchuan Zhang
摘要
Open-vocabulary panoptic segmentation aims to segment and classify everything in diverse scenes across an unbounded vocabulary. Existing methods typically employ two-stage or single-stage framework. The two-stage framework involves cropping the image multiple times using masks generated by a mask generator, followed by feature extraction, while the single-stage framework relies on a heavyweight mask decoder to make up for the lack of spatial position information through self-attention and cross-attention in multiple stacked Transformer blocks. Both methods incur substantial computational overhead, thereby hindering the efficiency of model inference. To fill the gap in efficiency, we propose EOV-Seg, a novel single-stage, shared, efficient, and spatialaware framework designed for open-vocabulary panoptic segmentation. Specifically, EOV-Seg innovates in two aspects. First, a Vocabulary-Aware Selection (VAS) module is proposed to improve the semantic comprehension of visual aggregated features and alleviate the feature interaction burden on the mask decoder. Second, we introduce a Two-way Dynamic Embedding Experts (TDEE), which efficiently utilizes the spatial awareness capabilities of ViT-based CLIP backbone. To the best of our knowledge, EOV-Seg is the first open-vocabulary panoptic segmentation framework towards efficiency, which runs faster and achieves competitive performance compared with state-of-the-art methods. Specifically, with COCO training only, EOV-Seg achieves 24.5 PQ, 32.1 mIoU, and 11.6 FPS on the ADE20K dataset and the inference time of EOV-Seg is 4-19 times faster than state-of-theart methods. Especially, equipped with ResNet50 backbone, EOV-Seg runs 23.8 FPS with only 71M parameters on a single RTX 3090 GPU. Code is available at https://github.com/ nhw649/EOV-Seg .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Cross-Modality Perturbation Synergy Attack for Person Re-identificationYunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo 等NeurIPS 2024 · 被引用 67 次
- ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIPXin Niu, Manqi Zhao, Dongsheng Jiang, Yingying Wu 等CVPR 2026 · 被引用 5 次
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPeng Ren, Tian Bai, Jing Sun, Fuming SunICCV 2025 · 被引用 4 次
- What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image SegmentationJianghang Lin, Yue Hu, Jiangtao Shen, Yunhang Shen 等ACM MM 2025 · 被引用 1 次
- MARIS: Marine Open-Vocabulary Instance SegmentationBingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao 等CVPR 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen 等NeurIPS 2023 · 被引用 285 次
- Decoupling Zero-Shot Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Dengxin DaiCVPR 2022 · 被引用 255 次
相关 Paper
- Open-Vocabulary Universal Image Segmentation with MaskCLIPZheng Ding, Jieke Wang, Zhuowen TuICML 2023 · 被引用 150 次
- Open-vocabulary Panoptic Segmentation with Embedding ModulationXi Chen, Shuang Li, Ser-Nam Lim, Antonio Torralba 等ICCV 2023 · 被引用 42 次
- Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion ModelsJiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon 等CVPR 2023
- Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic SegmentationNikolay Kormushev, Josip Saric, Matej KristanCVPR 2026 · 被引用 2 次
- MasQCLIP for Open-Vocabulary Universal Image SegmentationXin Xu, Tianyi Xiong, Zheng Ding, Zhuowen TuICCV 2023 · 被引用 57 次
