Comp-Attn: Present-and-Align Attention for Compositional Video Generation
Hongyu Zhang, Yufan Deng, Shenghai Yuan, Xuehan Hou, Yian Zhao, Peng Jin, Chang Liu, Jie Chen
Abstract
In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relations, where the interaction and spatial relationship between subjects are misaligned. Existing methods adopt techniques, such as inference-time latent optimization or layout control, which fail to address both issues simultaneously. To tackle these problems, we propose Comp-Attn, a composition-aware cross-attention variant that follows a ``Present-and-Align” paradigm: it decouples the two challenges by enforcing subject presence at the condition level and achieving relational alignment at the attention-distribution level. Specifically, 1) We introduce Subject-aware Condition Interpolation (SCI) to reinforce subject-specific conditions and ensure each subject's presence; 2) We propose Layout-forcing Attention Modulation (LAM), which dynamically enforces the attention distribution to align with the relational layout of multiple subjects. Comp-Attn can be seamlessly integrated into various T2V baselines in a training-free manner, boosting T2V-CompBench scores by 15.7% and 11.7% on Wan2.1-T2V-14B and Wan2.2-T2V-A14B with only a 5% increase in inference time. Meanwhile, it also achieves strong performance on VBench and T2I-CompBench, demonstrating its scalability in general T2V and compositional text-to-image (T2I) tasks. Code and models are available at: https://github.com/Hong-yu-Zhang/Comp-attn.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53784ab6-8a64-4d6a-8db7-d12b53a71692Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel et al.ICCV 2023 · 800 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
Related papers
- TS-Attn: Temporal-wise Separable Attention for Multi-Event Video GenerationHongyu Zhang, Yufan Deng, Zilin Pan, Peng-Tao Jiang et al.ICLR 2026 · 5 citations
- TTOM: Test-Time Optimization and Memorization for Compositional Video GenerationLeigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang et al.ICLR 2026 · 6 citations
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin et al.CVPR 2025
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao et al.NeurIPS 2024 · 62 citations
- MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention GuidanceNathan Sala, Ofir Abramovich, Ariel Shamir, Daniel Cohen-Or et al.SIGGRAPH 2026
