Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
Haotian Dong, Wenjing Wang, Chen Li, Jing LYU, Di Lin
摘要
Generating RGB-A videos, which include alpha channels for transparency, has wide applications. However, current methods often suffer from low quality due to confusion between RGB and alpha. In this paper, we address this problem by learning shiftable RGB‑A distributions. We adjust both the latent space and noise space, shifting the alpha distribution outward while preserving the RGB distribution, thereby enabling stable transparency generation without compromising RGB quality. Specifically, for the latent space, we propose a transparency‑aware bidirectional diffusion loss during VAE training, which shifts the RGB‑A distribution according to likelihood. For the noise space, we propose shifting the mean of diffusion noise sampling and applying a Gaussian ellipse mask to provide transparency guidance and controllability. Additionally, we construct a high‑quality RGB‑A video dataset. Compared to state‑of‑the‑art methods, our model excels in visual quality, naturalness, transparency rendering, inference convenience, and controllability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 被引用 4 次
- LayerT2V: A Unified Multi-Layer Video Generation FrameworkGuangzhao Li, Kangrui Cen, Baixuan Zhao, Yi Xin 等ICML 2026 · 被引用 2 次
- UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion PriorsHouyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu 等SIGGRAPH 2026
它引用的顶会 Paper24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
相关 Paper
- Transparent Image Layer Diffusion using Latent TransparencyLvmin Zhang, Maneesh AgrawalaSIGGRAPH 2024 · 被引用 42 次
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li 等ACM MM 2025 · 被引用 2 次
- Trans-Adapter: A Plug-And-Play Framework for Transparent Image InpaintingYuekun Dai, Haitian Li, Shangchen Zhou, Chen Change LoyICCV 2025 · 被引用 1 次
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
- RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image DecompositionBinhao Wang, Shihao Zhao, Bo Cheng, Qiuyu Ji 等ICML 2026
