The Functional Correspondence Problem
Zihang Lai, Senthil Purushwalkam, Abhinav Gupta
Abstract
The ability to find correspondences in visual data is the essence of most computer vision tasks. But what are the right correspondences? The task of visual correspondence is well defined for two different images of same object instance. In case of two images of objects belonging to same category, visual correspondence is reasonably well-defined in most cases. But what about correspondence between two objects of completely different category – e.g., a shoe and a bottle? Does there exist any correspondence? Inspired by humans’ ability to: (a) generalize beyond semantic categories and; (b) infer functional affordances, we introduce the problem of functional correspondences in this paper. Given images of two objects, we ask a simple question: what is the set of correspondences between these two images for a given task? For example, what are the correspondences between a bottle and shoe for the task of pounding or the task of pouring. We introduce a new dataset: FunKPoint that has ground truth correspondences for 10 tasks and 20 object categories. We also introduce a modular task-driven representation for attacking this problem and demonstrate that our learned representation is effective for this task. But most importantly, because our supervision signal is not bound by semantics, we show that our learned representation can generalize better on few-shot classification problem. We hope this paper will inspire our community to think beyond semantics and focus more on cross-category generalization and learning representations for robotics tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 940c9ce9-cf4f-4c07-81f6-333be4dc798aCited by top-tier papers8
- Human Hands as Probes for Interactive Object UnderstandingMohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh GuptaCVPR 2022 · 26 citations
- Latent Implicit Visual ReasoningKelvin Li, Chuyi Shang, Leonid Karlinsky, Rogério Feris et al.CVPR 2026 · 15 citations
- UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual CorrespondenceRuihai Wu, Haoran Lu, Yiyan Wang, Yubo Wang et al.CVPR 2024 · 14 citations
- IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor ScenesQi Li, Kaichun Mo, Yanchao Yang, Hang Zhao et al.ICLR 2022 · 9 citations
- Do It Yourself: Learning Semantic Correspondence from Pseudo-LabelsOlaf Dünkel, Thomas Wimmer, Christian Theobalt, Christian Rupprecht et al.ICCV 2025 · 4 citations
Builds on9
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 247 citations
Related papers
- DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single DemoJunzhe Zhu, Yuanchen Ju, Junyi Zhang, Muhan Wang et al.ICLR 2025
- Weakly-Supervised Learning of Dense Functional CorrespondencesStefan Stojanov, Linan Zhao, Yunzhi Zhang, Daniel L. K. Yamins et al.ICCV 2025 · 2 citations
- Open-Vocabulary 3D Affordance Understanding via Functional Text Enhancement and Multilevel Representation AlignmentLin Wu, Wei Wei, Peizhuo Yu, Jianglin LanACM MM 2025
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen et al.CVPR 2021
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras et al.ICCV 2019 · 668 citations
