OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad
Luyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang, Yue Huang, Kun Zhang
Abstract
Although foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks. Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a7e009a-44b0-4b74-91c5-6362f1da5ec7Cited by top-tier papers4
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-DeconstructionLuyao Tang, Kunze Huang, Chaoqi Chen, Yuxuan Yuan et al.ICCV 2025 · 3 citations
- SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage EfficiencyZeqing Wang, Kangye Ji, Di Wang, Haibin Zhang et al.AAAI 2026 · 2 citations
- Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in VideosSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
- PRISM: Progressive Robust Learning for Open-World Continual Category DiscoveryWei Feng, Sijin Zhou, Yiwen Jiang, Zongyuan GeICLR 2026
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Disentangled Prompt Representation for Domain GeneralizationDe Cheng, Zhipeng Xu, Xinyang Jiang, Nannan Wang et al.CVPR 2024
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang et al.NeurIPS 2025 · 3 citations
- Reconstruct and Match: Out-of-Distribution Robustness via Topological HomogeneityChaoqi Chen, Luyao Tang, Hui HuangNeurIPS 2024 · 2 citations
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 35 citations
- Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation ModelsAmir Mohammad Karimi-Mamaghan, Samuele Papa, Karl Henrik Johansson, Stefan Bauer et al.ICLR 2025
