Unified Personalized Understanding, Generating and Editing
Yu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan, Haoyu Zheng, Liang Liang, Wenqiao Zhang, Feifei Shao, Haoyuan Li, Wanggui He, Hao Jiang, Yueting Zhuang
Abstract
Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a one-size-fits-all''paradigm and struggle to model user-specific concepts (e.g., generate a photo of ) in a consistent and controllable manner. Existing personalization methods typically rely on external retrieval, which is inefficient and poorly integrated into unified multimodal pipelines. Recent personalized unified models introduce learnable soft prompts to encode concept information, yet they either couple understanding and generation or depend on complex multi-stage training, leading to cross-task interference and ultimately to fuzzy or misaligned personalized knowledge. We present OmniPersona, an end-to-end personalization framework for unified LMMs that, for the first time, integrates personalized understanding, generation, and image editing within a single architecture. OmniPersona introduces structurally decoupled concept tokens, allocating dedicated subspaces for different tasks to minimize interference, and incorporates an explicit knowledge replay mechanism that propagates personalized attribute knowledge across tasks, enabling consistent personalized behavior. To systematically evaluate unified personalization, we propose OmniPBench, extending the public UnifyBench concept set with personalized editing tasks and cross-task evaluation protocols integrating understanding, generation, and editing. Experimental results demonstrate that OmniPersona delivers competitive and robust performance across diverse personalization tasks. We hope OmniPersona will serve as a strong baseline and spur further research on controllable, unified personalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a64eb26-80c0-4965-81c4-d44c8f619fa7Cited by top-tier papers2
- Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow MatchingOnkar Susladkar, Tushar Prakash, Gayatri Deshmukh, Kiet Nguyen et al.ICML 2026
- Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography AnalysisTianwei Lin, Zhongwei Qiu, Jie Cao, Jiang Liu et al.ICML 2026
Builds on22
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- UniTok: a Unified Tokenizer for Visual Generation and UnderstandingChuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang et al.NeurIPS 2025 · 164 citations
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- InstantBooth: Personalized Text-to-Image Generation without Test-Time FinetuningJing Shi, Wei Xiong, Zhe Lin, Hyun Joon JungCVPR 2024 · 115 citations
Related papers
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen et al.NeurIPS 2025 · 61 citations
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou et al.CVPR 2026 · 36 citations
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationTsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang et al.CVPR 2026 · 4 citations
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng et al.ICLR 2024 · 46 citations
- MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal ModelsMingrui Wu, Hang Liu, Jiayi Ji, Xiaoshuai Sun et al.CVPR 2026 · 5 citations
