FALCON: Fast Visual Concept Learning by Integrating Images, Linguistic descriptions, and Conceptual Relations
Lingjie Mei, Jiayuan Mao, Ziqi Wang, Chuang Gan, Joshua B. Tenenbaum
Abstract
We present a meta-learning framework for learning new visual concepts quickly, from just one or a few examples, guided by multiple naturally occurring data streams: simultaneously looking at images, reading sentences that describe the objects in the scene, and interpreting supplemental sentences that relate the novel concept with other concepts. The learned concepts support downstream applications, such as answering questions by reasoning about unseen images. Our model, namely FALCON, represents individual visual concepts, such as colors and shapes, as axis-aligned boxes in a high-dimensional space (the "box embedding space"). Given an input image and its paired sentence, our model first resolves the referential expression in the sentence and associates the novel concept with particular objects in the scene. Next, our model interprets supplemental sentences to relate the novel concept with other known concepts, such as "X has property Y" or "X is a kind of Y". Finally, it infers an optimal box embedding for the novel concept that jointly 1) maximizes the likelihood of the observed instances in the image, and 2) satisfies the relationships between the novel concepts and the known ones. We demonstrate the effectiveness of our model on both synthetic and real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d07a048-26f1-4df1-98b2-762701bfeb87Cited by top-tier papers5
- NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AIHanchen Yang, Zishen Wan, Ritik Raj, Joongun Park et al.DAC 2025 · 8 citations
- Interactive Visual Task Learning for RobotsWeiwei Gu, Anant Sah, Nakul GopalanAAAI 2024 · 8 citations
- CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-DesignZishen Wan, Hanchen Yang, Ritik Raj, Che-Kai Liu et al.HPCA 2025 · 6 citations
- REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic IntelligenceZishen Wan, Che-Kai Liu, Jiayi Qian, Hanchen Yang et al.HPCA 2026 · 2 citations
- Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-EncodingAlessandro Achille, Greg Ver Steeg, Tian Yu Liu, Matthew Trager et al.CVPR 2024
Builds on1
Related papers
- Concept Learners for Few-Shot LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2021 · 4 citations
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 8 citations
- Identifying Interpretable Subspaces in Image RepresentationsNeha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz et al.ICML 2023 · 41 citations
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris et al.CVPR 2021
- Meta Compositional Referring Expression SegmentationLi Xu, Mark He Huang, Xindi Shang, Zehuan Yuan et al.CVPR 2023
