Referring to Any Person
Qing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren, Yuda Xiong, Yihao Chen, Qin Liu, Lei Zhang
Abstract
Humans are undoubtedly the most important participants in computer vision, and the ability to detect any individual given a natural language description, a task we define as referring to any person, holds substantial practical value. However, we find that existing models generally fail to achieve real-world usability, and current benchmarks are limited by their focus on one-to-one referring, that hinder progress in this area. In this work, we revisit this task from three critical perspectives: task definition, dataset design, and model architecture. We first identify five aspects of referable entities and three distinctive characteristics of this task. Next, we introduce HumanRef, a novel dataset designed to tackle these challenges and better reflect real-world applications. From a model design perspective, we integrate a multimodal large language model with an object detection framework, constructing a robust referring model named RexSeek. Experimental results reveal that state-of-the-art models, which perform well on commonly used benchmarks like RefCOCO/+/g, struggle with HumanRef due to their inability to detect multiple individuals. In contrast, RexSeek not only excels in human referring but also generalizes effectively to common object referring, making it broadly applicable across various perception tasks. Code is available at https://github.com/IDEA-Research/RexSeek
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2a2476e-a8bb-40c9-95e0-20ddef1b4d5eCited by top-tier papers7
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong et al.CVPR 2026 · 79 citations
- WeDetect: Fast Open-Vocabulary Object Detection as RetrievalShenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu et al.CVPR 2026 · 11 citations
- Gaze Target Estimation Anywhere with ConceptsXu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou et al.CVPR 2026 · 3 citations
- Grounding Everything in Tokens for Multimodal Large Language ModelsXiangxuan Ren, Zhongdao Wang, Liping Hou, Pin Tang et al.CVPR 2026 · 2 citations
- AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View VideosTeng Yan, Yihan Liu, Jiongxu Chen, Teng Wang et al.CVPR 2026 · 1 citation
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
Related papers
- RefCrowd: Grounding the Target in Crowd with Referring ExpressionsHeqian Qiu, Hongliang Li, Taijin Zhao, Lanxiao Wang et al.ACM MM 2022 · 10 citations
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 157 citations
- Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu et al.CVPR 2024
- Refer to Any Segmentation Mask Group with Vision-Language PromptsShengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu et al.ICCV 2025
- Referring Multi-Object TrackingDongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong et al.CVPR 2023
