SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation
Yuxuan Zhang, Yiren Song, Jiaming Liu, Rui Wang, Jinpeng Yu, Hao Tang, Huaxia Li, Xu Tang, Yao Hu, Han Pan, Zhongliang Jing
摘要
Recent advancements in subject-driven image generation have led to zero-shot generation, yet precise selection and focus on crucial subject representations remain challenging. Addressing this, we introduce the SSR-Encoder, a novel architecture designed for selectively capturing any subject from single or multiple reference images. It responds to various query modalities including text and masks, without necessitating test-time fine-tuning. The SSR-Encoder combines a Token-to-Patch Aligner that aligns query inputs with image patches and a Detail-Preserving Subject Encoder for extracting and preserving fine features of the subjects, thereby generating subject embeddings. These embeddings, used in conjunction with original text embeddings, condition the generation process. Characterized by its model generalizability and efficiency, the SSR-Encoder adapts to a range of custom models and control modules. Enhanced by the Embedding Consistency Regularization Loss for improved training, our extensive experiments demonstrate its effectiveness in versatile and high-quality image generation, indicating its broad applicability. Project page: ssr-encoder.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper68
- MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence GenerationYiren Song, Cheng Liu, Mike Zheng ShouCVPR 2026 · 被引用 46 次
- Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationQihan Huang, Siming Fu, Jinlong Liu, Hao Jiang 等AAAI 2025 · 被引用 43 次
- Stable-Hair: Real-World Hair Transfer via Diffusion ModelYuxuan Zhang, Qing Zhang, Yiren Song, Jichao Zhang 等AAAI 2025 · 被引用 37 次
- OminiControl: Minimal and Universal Control for Diffusion TransformerZhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue 等ICCV 2025 · 被引用 34 次
- OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization DataYiren Song, Cheng Liu, Mike Zheng ShouNeurIPS 2025 · 被引用 33 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Nested Attention: Semantic-aware Attention Values for Concept PersonalizationOr Patashnik, Rinon Gal, Daniil Ostashev, Sergey Tulyakov 等SIGGRAPH 2025 · 被引用 6 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten 等SIGGRAPH 2024 · 被引用 57 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image CustomizationYeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin 等AAAI 2025 · 被引用 6 次
