AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image Synthesis
Zhenbo Shi, Zhi Chen, Zhenbo Xu, Wei Yang, Liusheng Huang
Abstract
Image generation from text is a challenging and ill-posed task. Images generated from previous methods usually have low semantic consistency with texts and the achieved resolution is limited. To generate semantically consistent high-resolution images, we propose a novel method named AtHom, in which two attention modules are developed to extract the relationships from both independent modality and unified modality. The first is a novel Independent Modality Attention Module (IAM), which is presented to find out semantically important areas in generated images and to extract the informative context in texts. The second is a new module named Unified Semantic Space Attention Module (UAM), which is utilized to find out the relationships between extracted text context and essential areas in generated images. In particular, to bring the semantic features of texts and images closer in a unified semantic space, AtHom incorporates a homomorphic training mode by exploiting an extra discriminator to distinguish between two different modalities. Extensive experiments show that our AtHom surpasses previous methods by large margins.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4efa456b-4b87-40dc-9727-9b0146c98103Cited by top-tier papers2
- Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image GenerationGuojin Zhong, Jin Yuan, Pan Wang, Kailun Yang et al.ACM MM 2023 · 7 citations
- EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video ReconstructionChengjie Ge, Xueyang Fu, Peng He, Kunyu Wang et al.AAAI 2025 · 6 citations
Related papers
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang et al.ICCV 2019 · 80 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham et al.ICLR 2023 · 8 citations
- Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationXintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding et al.ACM MM 2022 · 17 citations
