Hyden: A Hybrid Dual-Path Encoder for Monocular Geometry of High-resolution Images
Zaiwei Zhang, Marc Mapeke, Wei Ye, Rakesh Ranjan, JQ Huang
Abstract
We present a hybrid dual-path vision encoder (Hyden) for high-resolution monocular depth, point map and surface normal estimation, surpassing state-of-the-art accuracy with a fraction of the inference cost. The architecture pairs a lowresolution Vision Transformer branch for global context with a full-resolution CNN branch for fine details, fusing features via a lightweight MLP before decoding. By exploiting the linear scaling of CNNs and constraining transformer computation to a fixed resolution, the model delivers fast inference even on multimegapixel inputs. To overcome the scarcity of high-quality high-resolution supervision, we introduce a self-distillation framework that generates pseudo-labels from existing models at both lower resolution full images and high-resolution crops-global labels preserve geometric accuracy, while local labels capture sharper details. To demonstrate the flexibility of our approach, we integrate Hyden and our self-distillation method into DepthAnything-v2 for depth estimation and MoGe2 for surface normal and metric point map prediction, achieving stateof-the-art results on high-resolution benchmarks with the lowest inference latency among competing methods. 2 RELATED WORK 2.1 ZERO-SHOT MONOCULAR GEOMETRY ESTIMATION Traditional monocular models Bhat et al. (2021); Eigen et al. (2014); Li et al. (2022); Eigen & Fergus (2015); Saxena et al. ( 2008 ) were trained on single datasets for specific domains (e.g., indoor or street-view) and generalized poorly due to limited diversity and fixed camera setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D ScansAinaz Eftekhar, Alexander Sax, Jitendra Malik, Amir ZamirICCV 2021 · 422 citations
Related papers
- Any Resolution Any Geometry: From Multi-View To Multi-PatchWenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi et al.CVPR 2026 · 2 citations
- GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor ScenesChaoqiang Zhao, Matteo Poggi, Fabio Tosi, Lei Zhou et al.ICCV 2023 · 27 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
- Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic ScenesJiquan Zhong, Xiaolin Huang, Xiao YuACM MM 2023 · 6 citations
- Lite-Mono: A Lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth EstimationNing Zhang, Francesco Nex, George Vosselman, Norman KerleCVPR 2023
