Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open World
Xudong Wang, Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han
Abstract
Universal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Cost-Effective Communication: An Auction-based Method for Language Agent InteractionYijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang et al.AAAI 2026 · 15 citations
- All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker AdaptationXudong Wang, Gan Li, Zhiyu Liu, Yao Wang et al.ICLR 2026 · 4 citations
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process ThinkingYuehao Huang, Liang Liu, Shuangming Lei, Yukai Ma et al.ACM MM 2025 · 1 citation
- Lifelong Language-Conditioned Robotic Manipulation LearningXudong Wang, Zebin Han, Zhiyu Liu, Gan Li et al.AAAI 2026
Related papers
- Multimodal Causal Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Shuaifeng Li et al.NeurIPS 2025 · 1 citation
- Image-Adaptive YOLO for Object Detection in Adverse Weather ConditionsWenyu Liu, Gaofeng Ren, Runsheng Yu, Shi Guo et al.AAAI 2022 · 556 citations
- Unified Interaction Consistency Learning for Single-Source Domain-Generalized Object Detection in Urban ScenePeng Zhang, Xiang Yuan, Gong ChengAAAI 2026
- Detecting Everything in the Open World: Towards Universal Object DetectionZhenyu Wang, Yali Li, Xi Chen, Ser-Nam Lim et al.CVPR 2023
- UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language ModelingYuchuan Li, Azadeh Motamedi, Hyock Ju Kwon, Chul B Park et al.CVPR 2026
