Mimir: Improving Video Diffusion Models for Precise Text Understanding
Shuai Tan, Biao Gong, Yutong Feng, Kecheng Zheng, Dandan Zheng, Shuwei Shi, Yujun Shen, Jingdong Chen, Ming Yang
摘要
A majestic eagle soars above a vast, snow-covered forest, its powerful wings cutting through the crisp winter air. Below, a dense canopy of evergreen trees is blanketed in pristine snow, creating a serene landscape. A woman with short, curly silverhair in a red dress stands in a dimly lit, futuristic setting, looking to her left. Neon lights and digital screens fill the scene. By the end, she gazes away from the viewer against a neon cityscape, conveying mystery and anticipation. The transformation from a flower bud to a fully blooming flower. A vast, golden desert stretches endlessly under a brilliant blue sky in the early morning light. As the sun sets, the sky transforms into a canvas of vibrant oranges. Finally, the night falls, revealing a breathtaking canopy of stars, with the Milky Way arching gracefully over the tranquil.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang 等CVPR 2026 · 被引用 9 次
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong 等ICCV 2025 · 被引用 5 次
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment DesignWen-Fan Wang, Ting-Ying Lee, Chien-Ting Lu, Che-Wei Hsu 等UIST 2025 · 被引用 4 次
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesShuai Tan, Bill Gong, Bin Ji, Ye PanICCV 2025 · 被引用 3 次
- Stylized-Face: A Million-Level Stylized Face Dataset for Face RecognitionZhengyuan Peng, Jianqing Xu, Yuge Huang, Jinkun Hao 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency ControlsMinghui Hu, Jianbin Zheng, Chuanxia Zheng, Chaoyue Wang 等CVPR 2024
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- Beyond Simple Edits: X-Planner for Complex Instruction-Based Image EditingChun-Hsiao Yeh, Yilin Wang, Nanxuan Zhao, Richard Zhang 等AAAI 2026
- TransPixeler: Advancing Text-to-Video Generation with TransparencyLuozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang 等CVPR 2025
- ImageSet2Text: Describing Sets of Images Through TextPiera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia 等AAAI 2026 · 被引用 1 次
