How Shift Equivariance Impacts Metric Learning for Instance Segmentation
Josef Lorenz Rumberger, Xiaoyan Yu, Peter Hirsch, Melanie Dohmen, Vanessa Emanuela Guarino, Ashkan Mokarian, Lisa Mais, Jan Funke, Dagmar Kainmueller
摘要
Metric learning has received conflicting assessments concerning its suitability for solving instance segmentation tasks. It has been dismissed as theoretically flawed due to the shift equivariance of the employed CNNs and their respective inability to distinguish same-looking objects. Yet it has been shown to yield state of the art results for a variety of tasks, and practical issues have mainly been reported in the context of tile-and-stitch approaches, where discontinuities at tile boundaries have been observed. To date, neither of the reported issues have undergone thorough formal analysis. In our work, we contribute a comprehensive formal analysis of the shift equivariance properties of encoder-decoder-style CNNs, which yields a clear picture of what can and cannot be achieved with metric learning in the face of same-looking objects. In particular, we prove that a standard encoder-decoder network that takes d-dimensional images as input, with l pooling layers and pooling factor f , has the capacity to distinguish at most f dl same-looking objects, and we show that this upper limit can be reached. Furthermore, we show that to avoid discontinuities in a tile-and-stitch approach, assuming standard batch size 1, it is necessary to employ valid convolutions in combination with a training output window size strictly greater than f l , while at test-time it is necessary to crop tiles to size n • f l before stitching, with n ≥ 1. We complement these theoretical findings by discussing a number of insightful special cases for which we show empirical results on synthetic and real data. Code: https://github.com/Kainmueller-Lab/ shift_equivariance_unet * equal contribution
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Image Priors Through Patch-Based Diffusion Models for Solving Inverse ProblemsJason Hu, Bowen Song, Xiaojian Xu, Liyue Shen 等NeurIPS 2024 · 被引用 32 次
- ShapeEmbed: a self-supervised learning framework for 2D contour quantificationAnna Foix Romero, Craig Russell, Alexander Krull, Virginie UhlmannNeurIPS 2025 · 被引用 8 次
- Tiling Artifacts and Trade-Offs of Feature Normalization in the Segmentation of Large Biological ImagesElena Buglakova, Anwai Archit, Edoardo D'Imprima, Julia Mahamid 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper3
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan 等ICLR 2021 · 被引用 32 次
- On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial LocationOsman Semih Kayhan, Jan C. van GemertCVPR 2020
- Instance Segmentation of Biological Images Using Harmonic EmbeddingsVictor Kulikov, Victor S. LempitskyCVPR 2020
相关 Paper
- Improving Equivariance in State-of-the-Art Supervised Depth and Normal PredictorsYuanyi Zhong, Anand Bhattad, Yu-Xiong Wang, David A. ForsythICCV 2023 · 被引用 3 次
- SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth EstimationLijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin 等CVPR 2020
- Scale-Equivariant Steerable NetworksIvan Sosnovik, Michal Szmaja, Arnold W. M. SmeuldersICLR 2020 · 被引用 169 次
- Cross-Image-Attention for Conditional Embeddings in Deep Metric LearningDmytro Kotovenko, Pingchuan Ma, Timo Milbich, Björn OmmerCVPR 2023
- Towards Interpretable Deep Metric Learning with Structural MatchingWenliang Zhao, Yongming Rao, Ziyi Wang, Jiwen Lu 等ICCV 2021 · 被引用 52 次
