AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image Synthesis
Zhenbo Shi, Zhi Chen, Zhenbo Xu, Wei Yang, Liusheng Huang
摘要
Image generation from text is a challenging and ill-posed task. Images generated from previous methods usually have low semantic consistency with texts and the achieved resolution is limited. To generate semantically consistent high-resolution images, we propose a novel method named AtHom, in which two attention modules are developed to extract the relationships from both independent modality and unified modality. The first is a novel Independent Modality Attention Module (IAM), which is presented to find out semantically important areas in generated images and to extract the informative context in texts. The second is a new module named Unified Semantic Space Attention Module (UAM), which is utilized to find out the relationships between extracted text context and essential areas in generated images. In particular, to bring the semantic features of texts and images closer in a unified semantic space, AtHom incorporates a homomorphic training mode by exploiting an extra discriminator to distinguish between two different modalities. Extensive experiments show that our AtHom surpasses previous methods by large margins.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image GenerationGuojin Zhong, Jin Yuan, Pan Wang, Kailun Yang 等ACM MM 2023 · 被引用 7 次
- EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video ReconstructionChengjie Ge, Xueyang Fu, Peng He, Kunyu Wang 等AAAI 2025 · 被引用 6 次
相关 Paper
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 等ICCV 2019 · 被引用 80 次
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing 等CVPR 2022 · 被引用 296 次
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 被引用 39 次
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham 等ICLR 2023 · 被引用 8 次
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding 等ACM MM 2022 · 被引用 17 次
