Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction
Shanshan Huang, Haoxuan Li, Chunyuan Zheng, Mingyuan Ge, Wei Gao, Lei Wang, Li Liu
摘要
Fashion image editing is a valuable tool for designers to convey their creative ideas by visualizing design concepts. With the recent advances in text editing methods, significant progress has been made in fashion image editing. However, they face two key challenges: spurious correlations in training data often induce changes in other areas when editing an area representing the intended editing concept, and these models typically lack the ability to edit multiple concepts simultaneously. To address the above challenges, we propose a novel Text-driven Fashion Image ediTing framework called T-FIT to mitigate the impact of spurious correlation by integrating counterfactual reasoning with compositional concept learning to precisely ensure compositional multiconcept fashion image editing relying solely on text descriptions. Specifically, T-FIT includes three key components. (i) Counterfactual abduction module, which learns an exogenous variable of the source image by a denoising U-Net model. (ii) Concept learning module, which identifies concepts in fashion image editing-such as clothing types and colors and projects a target concept into the space spanned from a series of textual prompts. (iii) Concept composition module, which enables simultaneous adjustments of multiple concepts by aggregating each concept's direction vector obtained from the concept learning module. Extensive experiments show that our method can achieve state-of-theart performance on various fashion image editing tasks, including single-concept editing (e.g., sleeve length, clothing type) and multi-concept editing (e.g., color & sleeve length).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Unveiling Extraneous Sampling Bias with Data Missing-Not-At-RandomChunyuan Zheng, Haocheng Yang, Haoxuan Li, Mengyue YangNeurIPS 2025 · 被引用 15 次
- Addressing Correlated Latent Exogenous Variables in Debiased Recommender SystemsShuqiang Zhang, Yuchao Zhang, Jinkun Chen, Haochen SuiKDD 2025 · 被引用 4 次
- Mitigating Data Imbalance in Time Series Classification Based on Counterfactual Minority Samples AugmentationLei Wang, Shanshan Huang, Chunyuan Zheng, Jun Liao 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan 等ICCV 2023 · 被引用 770 次
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 被引用 276 次
相关 Paper
- TexFit: Text-Driven Fashion Image Editing with Diffusion ModelsTongxin Wang, Mang YeAAAI 2024 · 被引用 20 次
- Edit Like A Designer: Modeling Design Workflows for Unaligned Fashion EditingQiyu Dai, Shuai Yang, Wenjing Wang, Wei Xiang 等ACM MM 2021 · 被引用 5 次
- FashionTailor: Controllable Clothing Editing for Human Images with Appearance PreservingJie Hou, Jianghong Ma, Xiangyu Mu, Haijun Zhang 等AAAI 2025 · 被引用 1 次
- Doubly Abductive Counterfactual Inference for Text-Based Image EditingXue Song, Jiequan Cui, Hanwang Zhang, Jingjing Chen 等CVPR 2024 · 被引用 7 次
- FEAT: Fashion Editing and Try-On from Any DesignSoye Kwon, Keonyoung Lee, Dahuin Jung, Jaekoo LeeCVPR 2026
