OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad
Luyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang, Yue Huang, Kun Zhang
摘要
Although foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks. Code
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-DeconstructionLuyao Tang, Kunze Huang, Chaoqi Chen, Yuxuan Yuan 等ICCV 2025 · 被引用 3 次
- SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage EfficiencyZeqing Wang, Kangye Ji, Di Wang, Haibin Zhang 等AAAI 2026 · 被引用 2 次
- Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in VideosSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan 等AAAI 2026
- PRISM: Progressive Robust Learning for Open-World Continual Category DiscoveryWei Feng, Sijin Zhou, Yiwen Jiang, Zongyuan GeICLR 2026
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Disentangled Prompt Representation for Domain GeneralizationDe Cheng, Zhipeng Xu, Xinyang Jiang, Nannan Wang 等CVPR 2024
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang 等NeurIPS 2025 · 被引用 3 次
- Reconstruct and Match: Out-of-Distribution Robustness via Topological HomogeneityChaoqi Chen, Luyao Tang, Hui HuangNeurIPS 2024 · 被引用 2 次
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
- Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation ModelsAmir Mohammad Karimi-Mamaghan, Samuele Papa, Karl Henrik Johansson, Stefan Bauer 等ICLR 2025
