Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open World
Xudong Wang, Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han
摘要
Universal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Cost-Effective Communication: An Auction-based Method for Language Agent InteractionYijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang 等AAAI 2026 · 被引用 15 次
- All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker AdaptationXudong Wang, Gan Li, Zhiyu Liu, Yao Wang 等ICLR 2026 · 被引用 4 次
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu 等AAAI 2026 · 被引用 2 次
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process ThinkingYuehao Huang, Liang Liu, Shuangming Lei, Yukai Ma 等ACM MM 2025 · 被引用 1 次
- Lifelong Language-Conditioned Robotic Manipulation LearningXudong Wang, Zebin Han, Zhiyu Liu, Gan Li 等AAAI 2026
相关 Paper
- Multimodal Causal Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Shuaifeng Li 等NeurIPS 2025 · 被引用 1 次
- Image-Adaptive YOLO for Object Detection in Adverse Weather ConditionsWenyu Liu, Gaofeng Ren, Runsheng Yu, Shi Guo 等AAAI 2022 · 被引用 556 次
- Unified Interaction Consistency Learning for Single-Source Domain-Generalized Object Detection in Urban ScenePeng Zhang, Xiang Yuan, Gong ChengAAAI 2026
- Detecting Everything in the Open World: Towards Universal Object DetectionZhenyu Wang, Yali Li, Xi Chen, Ser-Nam Lim 等CVPR 2023
- UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language ModelingYuchuan Li, Azadeh Motamedi, Hyock Ju Kwon, Chul B Park 等CVPR 2026
