Lune

ICML2026顶会

Boost the Identity-Preserving Embedding for Consistent Visual Generation

Zixun Xia, Kai Wang, Shuyu Guo, Boqian Li, jian Yang, Yaxing Wang

出版方
2026年份

摘要

Text-to-image models have advanced high-fidelity content generation, but their inability to maintain subject consistency hampers realistic applications. Existing training-based methods rely on heavy computation and large datasets; while training-free approaches demand excessive memory or complex auxiliary modules. In this paper, we first reveal a key property overlooked in prior works that the identity-relevant signals, termed Identity-Preserving Embeddings ( IPemb ), are implicitly encoded in textual embeddings of frame prompts. To address the consistent T2I generation with the IPemb embedding, we propose Boost Identity-Preserving Embedding ( BIPE ), a training-free yet plug-and-play framework that explicitly extracts and enhances the IPemb . Its core innovations are two complementary techniques: First, Adaptive Singular-Value Rescaling ( adaSVR ) applies singular-value decomposition to the joint embedding matrix of all frame prompts, amplifying identity-centric components while suppressing frame-specific noise. Second, Union Key ( UniK ) further reinforces consistency by aligning the T2I backbone’s image-text attention across the entire generation sequence. Experiments on the ConsiStory+ benchmark demonstrate BIPE outperforms existing methods in both qualitative and quantitative metrics. To address the gap in evaluating a broader range of scenarios with diversified prompt templates, we introduce a DiverStory benchmark to further confirm our scalability.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext cff2c130-29f6-405b-97d4-2b302d15cede

它引用的顶会 Paper33

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖