Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera
Mukai Yu, Mosam Dabhi, Liuyue Xie, Sebastian Scherer, László A. Jeni
Abstract
Modern perception increasingly relies on fisheye, panoramic, and other wide field-of-view (FoV) cameras, yet most pipelines still apply planar CNNs designed for pinhole imagery on 2D grids, where pixel-space neighborhoods misrepresent physical adjacency and models are sensitive to global rotations. Traditional spherical CNNs partially address this mismatch but require costly spherical harmonic transform that constrains resolution and efficiency. We present Unified Spherical Frontend (USF), a distortion-free lens-agnostic framework that transforms images from any calibrated camera onto the unit sphere via ray-direction correspondences, and performs spherical resampling, convolution, and pooling canonically in the spatial domain. USF is modular: projection, location sampling, value interpolation, and resolution control are fully decoupled. Its configurable distanceonly convolution kernels offer rotation-equivariance, mirroring translation-equivariance in planar CNNs while avoiding harmonic transforms entirely. We compare multiple standard planar backbones with their spherical counterparts across classification, detection, and segmentation tasks on synthetic (Spherical MNIST) and real-world (PANDORA, Stanford 2D-3D-S) datasets, and stress-test robustness to extreme lens distortions, varying FoV, and arbitrary rotations. USF scales efficiently to high-resolution spherical imagery and maintains less than 1% performance drop under random test-time rotations without training-time rotational augmentation, and enables zero-shot generalization to any unseen (wide-FoV) lenses with minimal performance degradation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e780967-9968-48d8-abc0-309acb95a3fbBuilds on4
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang et al.ICLR 2024 · 126 citations
- Scaling Spherical CNNsCarlos Esteves, Jean-Jacques E. Slotine, Ameesh MakadiaICML 2023 · 28 citations
- Scalable and Equivariant Spherical CNNs by Discrete-Continuous (DISCO) ConvolutionsJeremy Ocampo, Matthew A. Price, Jason D. McEwenICLR 2023 · 5 citations
Related papers
- UniK3D: Universal Camera Monocular 3D EstimationLuigi Piccinelli, Christos Sakaridis, Mattia Segù, Yung-Hsu Yang et al.CVPR 2025
- PDO-eS2CNNs: Partial Differential Operator Based Equivariant Spherical CNNsZhengyang Shen, Tiancheng Shen, Zhouchen Lin, Jinwen MaAAAI 2021 · 26 citations
- Spin-Weighted Spherical CNNsCarlos Esteves, Ameesh Makadia, Kostas DaniilidisNeurIPS 2020 · 81 citations
- OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionYuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang et al.CVPR 2022 · 79 citations
- FisheyeHDK: Hyperbolic Deformable Kernel Learning for Ultra-Wide Field-of-View Image RecognitionOla Ahmad, Freddy LécuéAAAI 2022 · 21 citations
