Towards Universal Perception through Language-Guided Open-World Object Detection
Zihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long, Ke Li, Xing Sun, Jiao Xie, Shaohui Lin
摘要
Open-vocabulary object detection seeks to recognize objects from arbitrary language inputs, extending detection beyond fixed training categories. While recent methods have made progress in detecting unseen categories, they typically require a set of predefined categories during the inference stage, hindering practical deployment in open-world scenarios. To overcome this crucial limitation, we propose UniPerception , a novel universal perception framework based on open-vocabulary object detection. It not only excels at open-vocabulary object detection but is also capable of generating labels for target objects in the absence of predefined vocabularies, and can be adapted to a broad range of vision-language tasks simply by modifying the language instructions. UniPerception seamlessly integrates three key innovations: 1) a robust visual detector trained on diverse data sources to capture rich and generalizable visual representations; 2) a language model with interleaved cross-modality fusion layers to interpret instructions and generate fine-grained responses conditioned on visual features; and 3) a tailored multi-stage training strategy that effectively bridges detection-specific learning with general vision-language understanding. We conduct extensive experiments on multiple benchmarks for open-vocabulary object detection (COCO, LVIS, ODinW), referring expression comprehension (RefCOCO/+/g, D3), and vision-language understanding (Flickr30k, VQAv2, GQA). The results show that UniPerception achieves strong open-world generalization and multi-modal understanding, outperforming the existing state-of-the-art methods and establishing itself as a unified, instruction-driven perception system.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionLihong Huang, Sheng-hua Zhong, Zhi Zhang, Yan LiuAAAI 2026
- Universal Instance Perception as Object Discovery and RetrievalBin Yan, Yi Jiang, Jiannan Wu, Dong Wang 等CVPR 2023
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language InterfaceHao Tang, Chen-Wei Xie, Haiyang Wang, Xiaoyi Bao 等NeurIPS 2025 · 被引用 30 次
- Training-Free Open-Ended Object Detection and Segmentation via Attention as PromptsZhiwei Lin, Yongtao Wang, Zhi TangNeurIPS 2024 · 被引用 27 次
- Detecting Everything in the Open World: Towards Universal Object DetectionZhenyu Wang, Yali Li, Xi Chen, Ser-Nam Lim 等CVPR 2023
