How Shift Equivariance Impacts Metric Learning for Instance Segmentation
Josef Lorenz Rumberger, Xiaoyan Yu, Peter Hirsch, Melanie Dohmen, Vanessa Emanuela Guarino, Ashkan Mokarian, Lisa Mais, Jan Funke, Dagmar Kainmueller
Abstract
Metric learning has received conflicting assessments concerning its suitability for solving instance segmentation tasks. It has been dismissed as theoretically flawed due to the shift equivariance of the employed CNNs and their respective inability to distinguish same-looking objects. Yet it has been shown to yield state of the art results for a variety of tasks, and practical issues have mainly been reported in the context of tile-and-stitch approaches, where discontinuities at tile boundaries have been observed. To date, neither of the reported issues have undergone thorough formal analysis. In our work, we contribute a comprehensive formal analysis of the shift equivariance properties of encoder-decoder-style CNNs, which yields a clear picture of what can and cannot be achieved with metric learning in the face of same-looking objects. In particular, we prove that a standard encoder-decoder network that takes d-dimensional images as input, with l pooling layers and pooling factor f , has the capacity to distinguish at most f dl same-looking objects, and we show that this upper limit can be reached. Furthermore, we show that to avoid discontinuities in a tile-and-stitch approach, assuming standard batch size 1, it is necessary to employ valid convolutions in combination with a training output window size strictly greater than f l , while at test-time it is necessary to crop tiles to size n • f l before stitching, with n ≥ 1. We complement these theoretical findings by discussing a number of insightful special cases for which we show empirical results on synthetic and real data. Code: https://github.com/Kainmueller-Lab/ shift_equivariance_unet * equal contribution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 374618eb-5911-4a47-abe1-e27caf92ed5eCited by top-tier papers3
- Learning Image Priors Through Patch-Based Diffusion Models for Solving Inverse ProblemsJason Hu, Bowen Song, Xiaojian Xu, Liyue Shen et al.NeurIPS 2024 · 32 citations
- ShapeEmbed: a self-supervised learning framework for 2D contour quantificationAnna Foix Romero, Craig Russell, Alexander Krull, Virginie UhlmannNeurIPS 2025 · 8 citations
- Tiling Artifacts and Trade-Offs of Feature Normalization in the Segmentation of Large Biological ImagesElena Buglakova, Anwai Archit, Edoardo D'Imprima, Julia Mahamid et al.ICCV 2025 · 3 citations
Builds on3
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan et al.ICLR 2021 · 32 citations
- On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial LocationOsman Semih Kayhan, Jan C. van GemertCVPR 2020
- Instance Segmentation of Biological Images Using Harmonic EmbeddingsVictor Kulikov, Victor S. LempitskyCVPR 2020
Related papers
- Improving Equivariance in State-of-the-Art Supervised Depth and Normal PredictorsYuanyi Zhong, Anand Bhattad, Yu-Xiong Wang, David A. ForsythICCV 2023 · 3 citations
- SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth EstimationLijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin et al.CVPR 2020
- Scale-Equivariant Steerable NetworksIvan Sosnovik, Michal Szmaja, Arnold W. M. SmeuldersICLR 2020 · 169 citations
- Cross-Image-Attention for Conditional Embeddings in Deep Metric LearningDmytro Kotovenko, Pingchuan Ma, Timo Milbich, Björn OmmerCVPR 2023
- Towards Interpretable Deep Metric Learning with Structural MatchingWenliang Zhao, Yongming Rao, Ziyi Wang, Jiwen Lu et al.ICCV 2021 · 52 citations
