Lune

ACM MM2025Top-tier venue

SSAIM: Not All Self-Attentions Contain Effective Spatial Structure in Diffusion Models for Text-to-Image Editing

Zhenbo Yu, Jimin Dai, Yingzhen Zhang, Jian Yang, Lei Luo

2025Year
3Citations
1Top-tier citations

Abstract

With the rapid progress of diffusion-based Text-to-Image Generation (TIG), Text-to-Image Editing (TIE) has become increasingly important for enabling controllable visual content creation. A core challenge in TIE is generating text-guided edits while preserving the spatial structure of the original image. Recent methods attempt to address this by leveraging self-attention maps from diffusion models, as these encode rich spatial information. However, we identify two key limitations: (1) not all self-attention maps contribute meaningfully to spatial structure, and (2) over-reliance on them can suppress desired editing effects. To address this, we propose the Spatial Information Score (SIS), a novel metric that quantifies the spatial structure encoded in each self-attention map. Leveraging SIS, we develop Selective Self-Attention-based Image Manipulation (SSAIM), which selectively utilizes self-attention maps with effective spatial structure (high SIS) to preserve the structural of the original image and reduce excessive reliance on self-attention maps with ineffective spatial structure (low SIS) to enhance editing performance in TIE tasks. Extensive experiments across diverse TIE tasks demonstrate that SSAIM significantly improves both structural fidelity and editing quality.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get b8e6e2ed-0048-4b09-a356-cd387e027344

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines