SSAIM: Not All Self-Attentions Contain Effective Spatial Structure in Diffusion Models for Text-to-Image Editing
Zhenbo Yu, Jimin Dai, Yingzhen Zhang, Jian Yang, Lei Luo
Abstract
With the rapid progress of diffusion-based Text-to-Image Generation (TIG), Text-to-Image Editing (TIE) has become increasingly important for enabling controllable visual content creation. A core challenge in TIE is generating text-guided edits while preserving the spatial structure of the original image. Recent methods attempt to address this by leveraging self-attention maps from diffusion models, as these encode rich spatial information. However, we identify two key limitations: (1) not all self-attention maps contribute meaningfully to spatial structure, and (2) over-reliance on them can suppress desired editing effects. To address this, we propose the Spatial Information Score (SIS), a novel metric that quantifies the spatial structure encoded in each self-attention map. Leveraging SIS, we develop Selective Self-Attention-based Image Manipulation (SSAIM), which selectively utilizes self-attention maps with effective spatial structure (high SIS) to preserve the structural of the original image and reduce excessive reliance on self-attention maps with ineffective spatial structure (low SIS) to enhance editing performance in TIE tasks. Extensive experiments across diverse TIE tasks demonstrate that SSAIM significantly improves both structural fidelity and editing quality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b8e6e2ed-0048-4b09-a356-cd387e027344Cited by top-tier papers1
Ask how each one uses itRelated papers
- Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image EditingBingyan Liu, Chengyu Wang, Tingfeng Cao, Kui Jia et al.CVPR 2024
- LUSD: Localized Update Score Distillation for Text-Guided Image EditingWorameth Chinchuthakun, Tossaporn Saengja, Nontawat Tritrong, Pitchaporn Rewatbowornwong et al.ICCV 2025
- Conditional Score Guidance for Text-Driven Image-to-Image TranslationHyunsoo Lee, Minsoo Kang, Bohyung HanNeurIPS 2023 · 23 citations
- LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image EditingAchint Soni, Meet Soni, Sirisha RambhatlaICCV 2025 · 1 citation
- SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image GenerationChengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai et al.AAAI 2024 · 30 citations
