MONET: Multiview Semi-Supervised Keypoint Detection via Epipolar Divergence
Yuan Yao, Yasamin Jafarian, Hyun Soo Park
Abstract
This paper presents MONET---an end-to-end semi-supervised learning framework for a keypoint detector using multiview image streams. In particular, we consider general subjects such as non-human species where attaining a large scale annotated dataset is challenging. While multiview geometry can be used to self-supervise the unlabeled data, integrating the geometry into learning a keypoint detector is challenging due to representation mismatch. We address this mismatch by formulating a new differentiable representation of the epipolar constraint called epipolar divergence---a generalized distance from the epipolar lines to the corresponding keypoint distribution. Epipolar divergence characterizes when two view keypoint distributions produce zero reprojection error. We design a twin network that minimizes the epipolar divergence through stereo rectification that can significantly alleviate computational complexity and sampling aliasing in training. We demonstrate that our framework can localize customized keypoints of diverse species, e.g., humans, dogs, and monkeys.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f55906c3-7156-49da-aba0-fb8ae8e703eeCited by top-tier papers7
- Dens3R: A Foundation Model for 3D Geometry PredictionXianze Fang, Jingnan Gao, Zhe Wang, Zhuo Chen et al.ICLR 2026 · 45 citations
- Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated ObjectsAtsuhiro Noguchi, Umar Iqbal, Jonathan Tremblay, Tatsuya Harada et al.CVPR 2022 · 32 citations
- Multiview Human Body Reconstruction from Uncalibrated CamerasZhixuan Yu, Linguang Zhang, Yuanlu Xu, Chengcheng Tang et al.NeurIPS 2022 · 24 citations
- PUMP: Pyramidal and Uniqueness Matching Priors for Unsupervised Learning of Local DescriptorsJérôme Revaud, Vincent Leroy, Philippe Weinzaepfel, Boris ChidlovskiiCVPR 2022 · 16 citations
- Epipolar TransformersYihui He, Rui Yan, Katerina Fragkiadaki, Shoou-I YuCVPR 2020
Related papers
- Dense Keypoints via Multiview SupervisionZhixuan Yu, Haozheng Yu, Long Sha, Sujoy Ganguly et al.NeurIPS 2021
- Semi-supervised Keypoint LocalizationOlga Moskvyak, Frédéric Maire, Feras Dayoub, Mahsa BaktashmotlaghICLR 2021 · 17 citations
- Triangulation Residual Loss for Data-efficient 3D Pose EstimationJiachen Zhao, Tao Yu, Liang An, Yipeng Huang et al.NeurIPS 2023 · 13 citations
- Self-Supervised Learning of Interpretable Keypoints From Unlabelled VideosTomas Jakab, Ankush Gupta, Hakan Bilen, Andrea VedaldiCVPR 2020
- Digging into Uncertainty in Self-supervised Multi-view StereoHongbin Xu, Zhipeng Zhou, Yali Wang, Wenxiong Kang et al.ICCV 2021 · 68 citations
