Re-Aligning Language to Visual Objects with an Agentic Workflow
Yuming Chen, Jiangyan Feng, Haodong Zhang, Lijun Gong, Feng Zhu, Rui Zhao, Qibin Hou, Ming-Ming Cheng, Yibing Song
摘要
Language-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During the training process, recent studies leverage vision-language models (VLMs) to automatically generate human-like expressions for visual objects, facilitating training data scaling up. In this process, we observe that VLM hallucinations bring inaccurate object descriptions (e.g., object name, color, and shape) to deteriorate VL alignment quality. To reduce VLM hallucinations, we propose an agentic workflow controlled by a large language model (LLM) to re-align language to visual objects via adaptively adjusting image and text prompts. We name this workflow Real-LOD, which includes planning, tool use, and reflection steps. Given an image with detected objects and VLM raw language expressions, Real-LOD reasons its state automatically and arranges action based on our neural symbolic designs (i.e., planning). The action will adaptively adjust the image and text prompts, and send them to VLMs for object re-description (i.e., tool use). Then, we use another LLM to analyze these refined expressions for feedback (i.e., reflection). These steps are conducted in a cyclic form to gradually improve language descriptions for re-aligning to visual objects. We construct a dataset that contains a tiny amount of 0.18M images with re-aligned language expression and train a prevalent LOD model to surpass existing LOD methods by around 50% on the standard benchmarks. With automatic VL refinement, our Real-LOD workflow reveals a potential to preserve data quality along with scaling up data quantity, further improving LOD performance from a data-alignment perspective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive AlignmentZiheng Ouyang, Yiren Song, Yaoli Liu, Shihao Zhu 等CVPR 2026 · 被引用 6 次
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingInha Kang, Youngsun Lim, Seonho Lee, Jiho Choi 等ICLR 2026 · 被引用 1 次
- From Text to Simulation: A Multi-Agent LLM Workflow for Automated Chemical Process DesignXufei Tian, Wenli Du, Shaoyi Yang, Han Hu 等AAAI 2026 · 被引用 1 次
- Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency LearningBozhao Li, Shaocong Wu, Tong Shao, Senqiao Yang 等CVPR 2026
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained DescriptorsSheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu 等ICLR 2024 · 被引用 48 次
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie 等CVPR 2025
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- InstructDET: Diversifying Referring Object Detection with Generalized InstructionsRonghao Dang, Jiangyan Feng, Haodong Zhang, Chongjian Ge 等ICLR 2024 · 被引用 16 次
