ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
Xujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie, Zequn Jie, Chengwen Huang, Jianqing Peng, Xiaodan Liang
Abstract
Cross-modal fashion image synthesis has emerged as one of the most promising directions in the generation domain due to the vast untapped potential of incorporating multiple modalities and the wide range of fashion image applications. To facilitate accurate generation, cross-modal synthesis methods typically rely on Contrastive Language-Image Pre-training (CLIP) to align textual and garment information. In this work, we argue that simply aligning texture and garment information is not sufficient to capture the semantics of the visual information and therefore propose MaskCLIP. MaskCLIP decomposes the garments into semantic parts, ensuring fine-grained and semantically accurate alignment between the visual and text information. Building on MaskCLIP, we propose ARMANI, a unified cross-modal fashion designer with part-level garment-text alignment. ARMANI discretizes an image into uniform tokens based on a learned cross-modal codebook in its first stage and uses a Transformer to model the distribution of image tokens for a real image given the tokens of the control signals in its second stage. Contrary to prior approaches that also rely on two-stage paradigms, ARMANI introduces textual tokens into the codebook, making it possible for the model to utilize fine-grain semantic information to generate more realistic images. Further, by introducing a cross-modal Transformer, ARMANI is versatile and can accomplish image synthesis from various control signals, such as pure text, sketch images, and partial images. Extensive experiments conducted on our newly collected cross-modal fashion dataset demonstrate that ARMANI generates photo-realistic images in diverse synthesis tasks and outperforms existing state-of-the-art cross-modal image synthesis approaches. Our code is available at https://github.com/Harvey594/ARMANI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- SGDiff: A Style Guided Diffusion Model for Fashion SynthesisZhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. MokACM MM 2023 · 44 citations
- DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic AlignmentXujie Zhang, Binbin Yang, Michael C. Kampffmeyer, Wenqing Zhang et al.ICCV 2023 · 23 citations
- Tell2Design: A Dataset for Language-Guided Floor Plan GenerationSicong Leng, Yang Zhou, Mohammed Haroon Dupty, Wee Sun Lee et al.ACL 2023 · 20 citations
- ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World DataMaya Varma, Jean-Benoit Delbrouck, Sarah M. Hooper, Akshay Chaudhari et al.ICCV 2023 · 15 citations
- HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion ModelsZhifeng Xie, Hao Li, Huiming Ding, Mengtian Li et al.AAAI 2025 · 12 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
Related papers
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 33 citations
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.CVPR 2022 · 94 citations
- Towards Counterfactual Image Manipulation via CLIPYingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang et al.ACM MM 2022 · 33 citations
- HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImageJunyi Guo, Jingxuan Zhang, Fangyu Wu, Huanda Lu et al.ICCV 2025 · 1 citation
- TANGO: Text-driven Photorealistic and Robust 3D Stylization via Lighting DecompositionYongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang et al.NeurIPS 2022 · 112 citations
