Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Gangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang, Jingfeng Yao, Lianghui Zhu, Yuechuan Pu, Cheng Chi, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye
摘要
This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation models fine-tune Stable Diffusion and achieve impressive performance. However, they require a VAE to compress depth maps into latent space, which inevitably introduces flying pixels at edges and details. Our model addresses this challenge by directly performing diffusion generation in the pixel space, avoiding VAE-induced artifacts. To overcome the high complexity associated with pixel-space generation, we introduce two novel designs: 1) Semantics-Prompted Diffusion Transformers (SP-DiT), which incorporate semantic representations from vision foundation models into DiT to prompt the diffusion process, thereby preserving global semantic consistency while enhancing fine-grained visual details; and 2) Cascade DiT Design that progressively increases the number of tokens to further enhance efficiency and accuracy. Our model achieves the best performance among all published generative models across five benchmarks, and significantly outperforms all other models in edge-aware point cloud evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsRuicheng Wang, Sicheng Xu, Yue Dong, Yu Deng 等NeurIPS 2025 · 被引用 308 次
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng 等CVPR 2026 · 被引用 42 次
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie 等NeurIPS 2025 · 被引用 26 次
- InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsHao Yu, Haotong Lin, Jiawei Wang, Jiaxin Li 等CVPR 2026 · 被引用 19 次
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等ICML 2026 · 被引用 12 次
它引用的顶会 Paper53
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng 等CVPR 2026 · 被引用 82 次
- JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion TransformersByung-Ki Kwon, Qi Dai, Lee Hyoseok, Chong Luo 等ICCV 2025 · 被引用 3 次
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang 等NeurIPS 2025 · 被引用 18 次
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- FiffDepth: Feed-Forward Transformation of Diffusion-Based Generators for Detailed Depth EstimationYunpeng Bai, Qixing HuangICCV 2025 · 被引用 5 次
