Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
Yiqing Shi, Yiren Song, Mike Zheng Shou
摘要
Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm and show that image editing diffusion models are inherently image-to-image consistent, providing a more suitable foundation for dense perception task. We introduce Edit2Perceive, a unified diffusion framework that adapts editing models for depth, normal, and matting. Built upon the FLUX.1 Kontext architecture, our approach employs full-parameter fine-tuning and a pixel-space consistency loss to enforce structure-preserving refinement across intermediate denoising states. Moreover, our single-step deterministic inference yields up to faster runtime while training on relatively small datasets. Extensive experiments demonstrate comprehensive state-of-the-art results across all three tasks, revealing the strong potential of editing-oriented diffusion transformers for geometry-aware perception.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
相关 Paper
- FE2E: From Editor to Dense Geometry EstimatorJiyuan Wang, Chunyu Lin, Lei Sun, Rongying Liu 等CVPR 2026
- What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan 等ICLR 2025
- Image Editing As Programs with Diffusion ModelsYujia Hu, Songhua Liu, Zhenxiong Tan, Xingyi Yang 等NeurIPS 2025 · 被引用 10 次
- Training-Free Text-Guided Color Editing with Multi-Modal Diffusion TransformerZixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou 等ICLR 2026 · 被引用 6 次
- ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene EditingJun-Kun Chen, Samuel Rota Bulò, Norman Müller, Lorenzo Porzi 等CVPR 2024
