Yo'Chameleon: Personalized Vision and Language Generation
Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui, Yong Jae Lee, Yuheng Li
2025Year
7Top-tier citations
Abstract
https://thaoshibe.github.io/YoChameleon Figure 1 . Using only 3-5 images of a novel concept/subject, we personalize Large Multimodal Models (e.g., Chameleon [1]) so that they retain their original capabilities while enabling tailored language and vision generation for the novel concept.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bd6efeb-28a2-46e3-944a-3a2ea17a569dCited by top-tier papers7
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen et al.NeurIPS 2025 · 61 citations
- Unified Personalized Understanding, Generating and EditingYu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan et al.CVPR 2026 · 7 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
- Multimodal Policy Internalization for Conversational AgentsZhenhailong Wang, Jiateng Liu, Amin Fazel, Ritesh Sarkhel et al.ICLR 2026 · 2 citations
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image GenerationFangtai Wu, Mushui Liu, Weijie He, Zhao Wang et al.CVPR 2026 · 1 citation
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Yo'LLaVA: Your Personalized Language and Vision AssistantThao Nguyen, Haotian Liu, Yuheng Li, Mu Cai et al.NeurIPS 2024 · 68 citations
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and GenerationChengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma et al.CVPR 2025
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
- ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language TasksAndrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte et al.CVPR 2024 · 1 citation
- Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and BeyondYongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie et al.ACL 2024 · 9 citations
