PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer
Abstract
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions (e.g., MoGe), or necessitate compressing geometry into latent spaces (e.g., GeometryCrafter) to leverage pre-trained latent diffusion models. In this work, we demonstrate that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer built on a plain ViT, which operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion-based approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. We show that this streamlined approach yields results superior to complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, our model produces sharper geometric structures and achieves significantly better results on highly ambiguous regions, such as transparent objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4500278-0c62-4704-bbb4-eefffa6c20efBuilds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- PixNerd: Pixel Neural Field DiffusionShuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang et al.ICLR 2026 · 78 citations
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng et al.CVPR 2026 · 82 citations
- HORT: Monocular Hand-held Objects Reconstruction with TransformersZerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia SchmidICCV 2025 · 4 citations
- Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image GenerationAlan Baade, Eric Chan, Kyle Sargent, Changan Chen et al.ICML 2026 · 25 citations
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single ImagesPhilipp Wulff, Felix Wimbauer, Dominik Muhle, Daniel CremersICCV 2025 · 1 citation
