MIST: Multiple Instance Spatial Transformer
Baptiste Angles, Yuhe Jin, Simon Kornblith, Andrea Tagliasacchi, Kwang Moo Yi
Abstract
We propose a deep network that can be trained to tackle image reconstruction and classification problems that involve detection of multiple object instances, without any supervision regarding their whereabouts. The network learns to extract the most significant K patches, and feeds these patches to a task-specific network -e.g., auto-encoder or classifier -to solve a domain specific problem. The challenge in training such a network is the non-differentiable top-K selection process. To address this issue, we lift the training optimization problem by treating the result of top-K selection as a slack variable, resulting in a simple, yet effective, multi-stage training. Our method is able to learn to detect recurring structures in the training dataset by learning to reconstruct images. It can also learn to localize structures when only knowledge on the occurrence of the object is provided, and in doing so it outperforms the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f123ad4-fa88-4a6a-b166-3f047f5c10c9Cited by top-tier papers3
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 43 citations
- TUSK: Task-Agnostic Unsupervised KeypointsYuhe Jin, Weiwei Sun, Jan Hosang, Eduard Trulls et al.NeurIPS 2022 · 6 citations
- Movies2Scenes: Using Movie Metadata to Learn Scene RepresentationShixing Chen, Chun-Hao Liu, Xiang Hao, Xiaohan Nie et al.CVPR 2023
Builds on3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Weakly Supervised Object Detection With Segmentation CollaborationXiaoyan Li, Meina Kan, Shiguang Shan, Xilin ChenICCV 2019 · 105 citations
- Linearized Multi-Sampling for Differentiable Image TransformationWei Jiang, Weiwei Sun, Andrea Tagliasacchi, Eduard Trulls et al.ICCV 2019 · 24 citations
Related papers
- Robust Instance Segmentation Through Reasoning About Multi-Object OcclusionXiaoding Yuan, Adam Kortylewski, Yihong Sun, Alan L. YuilleCVPR 2021
- Estimating Low-Rank Region Likelihood MapsGabriela Csurka, Zoltan Kato, Andor Juhasz, Martin HumenbergerCVPR 2020
- Differentiable Patch Selection for Image RecognitionJean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn et al.CVPR 2021
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
- AutoGPart: Intermediate Supervision Search for Generalizable 3D Part SegmentationXueyi Liu, Xiaomeng Xu, Anyi Rao, Chuang Gan et al.CVPR 2022 · 17 citations
