π3: Permutation-Equivariant Visual Geometry Learning
Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, Tong He
Abstract
We introduce π 3 , a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, π 3 employs a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design not only makes our model inherently robust to input ordering, but also leads to higher accuracy and performance. These advantages enable our simple and bias-free approach to achieve state-ofthe-art performance on a wide range of tasks, including camera pose estimation, monocular/video depth estimation, and dense point map reconstruction. Code and models are available at Pi3. INTRODUCTION Visual geometry reconstruction, a long-standing and fundamental problem in computer vision, holds substantial potential for applications such as augmented reality (Engel et al., 2023 ), robotics (Zhu et al., 2024), and autonomous navigation (Mur-Artal et al., 2015) . While traditional methods addressed this challenge using iterative optimization techniques like Bundle Adjustment (BA) (Hartley & Zisserman, 2003) , the field has recently seen remarkable progress with feed-forward neural networks. End-to-end models like DUSt3R (Wang et al., 2024) and its successors have demonstrated the power of deep learning for reconstructing geometry from image pairs (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 283c51c9-2ee3-4e67-9e4f-86f47d13dfb8Cited by top-tier papers66
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- V-DPM: 4D Video Reconstruction with Dynamic Point MapsEdgar Sucar, Eldar Insafutdinov, Zihang Lai, Andrea VedaldiCVPR 2026 · 29 citations
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- Block-Sparse Global Attention for Efficient Multi-View Geometry TransformersChung-Shien Brian Wang, Christian Schmidt, Jens Piekenbrinck, Bastian LeibeCVPR 2026 · 22 citations
Builds on21
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
Related papers
- VGGT: Visual Geometry Grounded TransformerJianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi et al.CVPR 2025
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionWenyu Li, Sidun Liu, Peng Qiao, Yong DouACM MM 2025 · 3 citations
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionRuicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang et al.CVPR 2025
