MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
Run Ling, Ke Cao, Jian Lu, Ao Ma, Haowei Liu, Runze He, Changwei Wang, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Peng Wu, Guibing Guo
摘要
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster GenerationYuxin Qin, Ke Cao, Haowei Liu, Ao Ma 等CVPR 2026 · 被引用 5 次
- HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product ImagesYi Chen Liu, Donghao Zhou, Jie Wang, Xin Gao 等CVPR 2026 · 被引用 5 次
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationDonghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen 等ICCV 2025 · 被引用 128 次
- WISA: World simulator assistant for physics-aware text-to-video generationJing Wang, Ao Ma, Ke Cao, Jun Zheng 等NeurIPS 2025 · 被引用 93 次
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang 等ICCV 2025 · 被引用 58 次
相关 Paper
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject DisentanglementYufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang 等ICLR 2026 · 被引用 20 次
- BindWeave: Subject-Consistent Video Generation via Cross-Modal IntegrationZhaoyang Li, Dongjun Qian, Kai Su, qishuai diao 等ICLR 2026 · 被引用 23 次
- Human-Centric Video Generation via Collaborative Multi-Modal ConditioningLiyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li 等AAAI 2026 · 被引用 1 次
- MV-S2V: Multi-View Subject-Consistent Video GenerationZiyang Song, Xinyu Gong, Bangya Liu, Zelin ZhaoSIGGRAPH 2026 · 被引用 1 次
- UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image GenerationDanning Zhang, Yijing Lin, Shuhan Zhuang, Mengqi Huang 等ICML 2026
