CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer, Efstratios Gavves, Aishwarya Agrawal
Abstract
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains, including complex real-world scenes. However, these models suffer from a key limitation: they lack controllability. Specifically, current object-centric models learn representations based on their preconceived understanding of objects, without allowing user input to guide which objects are represented. Introducing controllability into object-centric models could unlock a range of useful capabilities, such as the ability to extract instance-specific representations from a scene. In this work, we propose a novel approach for user-directed control over slot representations by conditioning slots on language descriptions. The proposed CONTROLLABLE OBJECT-CENTRIC REPRESENTATION LEARNING approach, which we term CTRL-O, achieves targeted object-language binding in complex real-world scenes without requiring mask supervision. Next, we apply these controllable slot representations on two downstream vision language tasks: textto-image generation and visual question answering. The proposed approach enables instance-specific text-to-image generation and also achieves strong performance on visual question answering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0553dde2-a7fb-4e35-b897-377ca5a73638Cited by top-tier papers7
- Object-Centric Concept-BottlenecksDavid Steinmann, Wolfgang Stammer, Antonia Wüst, Kristian KerstingNeurIPS 2025 · 12 citations
- Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMsTiancheng Gu, Kaicheng Yang, Ziyong Feng, Xingjun Wang et al.ACM MM 2025 · 6 citations
- FORLA: Federated Object-Centric Representation Learning with Slot AttentionGuiqiu Liao, Matjaz Jogan, Eric Eaton, Daniel A. HashimotoNeurIPS 2025 · 3 citations
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 2 citations
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius et al.CVPR 2025
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- GLASS: Guided Latent Slot Diffusion for Object-Centric LearningKrishnakant Singh, Simone Schaub-Meyer, Stefan RothCVPR 2025
- Cycle Consistency Driven Object DiscoveryAniket Rajiv Didolkar, Anirudh Goyal, Yoshua BengioICLR 2024 · 10 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 29 citations
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
