Unified Dense Prediction of Video Diffusion
Lehan Yang, Lu Qi, Xiangtai Li, Sheng Li, Varun Jampani, Ming-Hsuan Yang
摘要
We present a unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts. We utilize colormap to represent entity masks and depth maps, tightly integrating dense prediction with RGB video generation. Introducing dense prediction information improves video generation's consistency and motion smoothness without increasing computational costs. Incorporating learnable task embeddings brings multiple dense prediction tasks into a single model, enhancing flexibility and further boosting performance. We further propose a large-scale dense prediction video dataset Panda-Dense, addressing the issue that existing datasets do not concurrently contain captions, videos, segmentation, or depth maps. Comprehensive experiments demonstrate the high efficiency of our method, surpassing the state-of-the-art in terms of video quality, consistency, and motion smoothness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu 等CVPR 2026 · 被引用 26 次
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 等AAAI 2026 · 被引用 14 次
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video GenerationHao Zhang, Chun-Han Yao, Simon Donné, Narendra Ahuja 等NeurIPS 2025 · 被引用 8 次
- DGS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View ReconstructionMeixi Song, Xin Lin, Dizhe Zhang, Haodong Li 等ICLR 2026 · 被引用 5 次
- MATRIX: Mask Track Alignment for Interaction-aware Video GenerationSiyoon Jin, Seongchan Kim, Jae Ho Lee, Dahyun Chung 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper35
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video GenerationJiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 等CVPR 2026 · 被引用 12 次
- High Quality Entity SegmentationLu Qi, Jason Kuen, Tiancheng Shen, Jiuxiang Gu 等ICCV 2023 · 被引用 91 次
- UniGS: Unified Representation for Image Generation and SegmentationLu Qi, Lehan Yang, Weidong Guo, Yu Xu 等CVPR 2024 · 被引用 4 次
- UniVS: Unified and Universal Video Segmentation with Prompts as QueriesMinghan Li, Shuai Li, Xindong Zhang, Lei ZhangCVPR 2024
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang 等ACM MM 2022 · 被引用 12 次
