HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver
Cong Wei, Yujie Zhong, Haoxian Tan, Yong Liu, Jie Hu, Dengjie Li, Zheng Zhao, Yujiu Yang
摘要
This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as well as the complex reasoning segmentation, make it difficult for them to handle various challenging instructions and achieve an accurate understanding of fine-grained vision-language correlations. We propose HyperSeg, the first VLLM-based universal segmentation model for pixel-level image and video perception, en compassing generic segmentation tasks and more complex reasoning perception tasks requiring powerful reasoning abilities and world knowledge. Besides, to fully leverage the recognition capabilities of VLLMs and the fine-grained visual information, HyperSeg incorporates hybrid entity recognition and fine-grained visual perceiver modules for various segmentation tasks. Combined with the temporal adapter, HyperSeg achieves a comprehensive understanding of temporal information. Experimental results validate the effectiveness of our insights in resolving universal image and video segmentation tasks, including the more complex reasoning perception tasks. Our code is available at https://github.com/congvvc/HyperSeg.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- UGround: Towards Unified Visual Grounding with Unrolled TransformersRui Qian, Xin Yin, Chuanhang Deng, Zhiyuan Peng 等ICML 2026 · 被引用 14 次
- Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object SegmentationHaichao Jiang, Tianming Liang, Wei-Shi Zheng, Jian-Fang HuCVPR 2026 · 被引用 7 次
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li 等ICML 2026 · 被引用 3 次
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 被引用 2 次
- WOW-Seg: A Word-free Open World Segmentation ModelDanyang Li, Tianhao Wu, Bin Lin, Zhenyuan Chen 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language ModelsCong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng 等ICCV 2025 · 被引用 8 次
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan 等NeurIPS 2024 · 被引用 186 次
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang 等NeurIPS 2024 · 被引用 147 次
- From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationYizhou Wang, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen 等ACL 2026
- ARGenSeg: Image Segmentation with Autoregressive Image Generation ModelXiaolong Wang, Lixiang Ru, Ziyuan Huang, Kaixiang Ji 等NeurIPS 2025 · 被引用 8 次
