Scaling Properties of Diffusion Models For Perceptual Tasks
Rahul Ravishankar, Zeeshan Patel, Jathushan Rajasegaran, Jitendra Malik
2025年份
6顶会引用
摘要
We fine-tune a pre-trained Diffusion Model (DM) for visual perception tasks. We take a RGB image, and a conditional image (i.e. next video frame, occlusion mask, etc.), along with the noised image of the ground truth prediction. Our model generates predictions for visual tasks such as depth estimation, optical flow prediction, and amodal segmentation, based on the conditional task embedding. We train a generalist model that can perform all three tasks with exceptional performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The Serial Scaling HypothesisYuxi Liu, Konpat Preechakul, Kananart Kuwaranancharoen, Yutong BaiICLR 2026 · 被引用 12 次
- LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video DiffusionFangfu Liu, Hao Li, Jiawei Chi, Hanyang Wang 等ICCV 2025 · 被引用 7 次
- gen2seg: Generative Models Enable Generalizable Instance SegmentationOm Khangaonkar, Hamed PirsiavashICLR 2026 · 被引用 2 次
- CondDiff-AMO: Integrating Conditional Diffusion Mechanism for Unified Amodal Mask GenerationCaijie Zhao, Bob ZhangAAAI 2026
- Motion Attribution for Video GenerationXindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba 等ICML 2026
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng 等NeurIPS 2025 · 被引用 45 次
- Visual Bridge: Universal Visual Perception Representations GeneratingYilin Gao, Shuguang Dou, Junzhou Li, Zhiheng Yu 等AAAI 2026 · 被引用 1 次
- Unleashing Text-to-Image Diffusion Models for Visual PerceptionWenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu 等ICCV 2023 · 被引用 327 次
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 等ICCV 2025 · 被引用 2 次
- What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan 等ICLR 2025
