YOLOE: Real-Time Seeing Anythi
Ao Wang, Lihao Liu, Hui Chen, Zijia Lin, Jungong Han, Guiguang Ding
摘要
Object detection and segmentation are widely employed in computer vision applications, yet conventional models like YOLO series, while efficient and accurate, are limited by predefined categories, hindering adaptability in open scenarios. Recent open-set methods leverage text prompts, visual cues, or prompt-free paradigm to overcome this, but often compromise between performance and efficiency due to high computational demands or deployment complexity. In this work, we introduce YOLOE, which integrates detection and segmentation across diverse open prompt mechanisms within a single highly efficient model, achieving real-time seeing anything. For text prompts, we propose Re-parameterizable Region-Text Alignment (RepRTA) strategy. It refines pretrained textual embeddings via a re-parameterizable lightweight auxiliary network and enhances visual-textual alignment with zero inference and transferring overhead. For visual prompts, we present Semantic-Activated Visual Prompt Encoder (SAVPE). It employs decoupled semantic and activation branches to bring improved visual embedding and accuracy with minimal complexity. For prompt-free scenario, we introduce Lazy Region-Prompt Contrast (LRPC) strategy. It utilizes a builtin large vocabulary and specialized embedding to identify all objects, avoiding costly language model dependency. Extensive experiments show YOLOE's exceptional zero-shot performance and transferability with high inference efficiency and low training cost. Notably, on LVIS, with less training cost and inference speedup, YOLOE-v8-S surpasses YOLO-Worldv2-S by 3.5 AP. When transferring to COCO, YOLOE-v8-L achieves and gains over closed-set YOLOv8-L with nearly less training time. Code and models are available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong 等CVPR 2026 · 被引用 79 次
- NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object DetectionYupeng Zhang, Ruize Han, Zhiwei Chen, Wei Feng 等CVPR 2026 · 被引用 2 次
- Lattice Boltzmann Model for Learning Real-World Pixel DynamicityGuangze Zheng, Shijie Lin, Haobo Zuo, Si Si 等NeurIPS 2025
- DOMR: Establishing Cross-View Segmentation via Dense Object MatchingJitong Liao, Yulu Gao, Shaofei Huang, Jialin Gao 等ACM MM 2025
- UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual PromptingGeonuk Kim, Minhoi Kim, Kangil Lee, Minsu Kim 等CVPR 2026
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
相关 Paper
- YOLO-World: Real-Time Open-Vocabulary Object DetectionTianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu 等CVPR 2024
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao 等CVPR 2024 · 被引用 8 次
- PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched TrainingWeifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li 等CVPR 2026 · 被引用 3 次
- CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingXiaoshi Wu, Feng Zhu, Rui Zhao, Hongsheng LiCVPR 2023
- Training-Free Open-Ended Object Detection and Segmentation via Attention as PromptsZhiwei Lin, Yongtao Wang, Zhi TangNeurIPS 2024 · 被引用 27 次
