Compass Control: Multi Object Orientation Control for Text-to-Image Generation
Rishubh Parihar, Vaibhav Agrawal, Sachidanand VS, Venkatesh Babu Radhakrishnan
Abstract
Personalization with 3D Orientation Control Few unposed Input Images 'A photo of V* car in front of the leaning tower of Pisa in Italy' 0.523 1.047 2.617 3.665 3.141 0.0 1.047 2.094 5.235 3.123 'A photo of a mother walking with a pram on a snowy street, festive Christmas lights, beautiful winter evening scene' 0.675, 0.725 0.80, 0.60 * equal contribution. † work done during an internship at VAL, IISc tion. In this work, we address the problem of multi-object orientation control in text-to-image diffusion models. This enables the generation of diverse multi-object scenes with precise orientation control for each object. The key idea is to condition the diffusion model with a set of orientationaware compass tokens, one for each object, along with text tokens. A light-weight encoder network predicts these com-This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f096378-f4e5-46f5-a38e-d6e1890be8eeCited by top-tier papers6
- Kontinuous Kontext: Continuous Strength Control for Instruction-based Image EditingRishubh Parihar, Or Patashnik, Daniil Ostashev, Venkatesh Babu Radhakrishnan et al.CVPR 2026 · 15 citations
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose ManipulationZhenyuan Qin, Xincheng Shuai, Henghui DingNeurIPS 2025 · 11 citations
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image GenerationVaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla et al.CVPR 2026 · 5 citations
- Zero-Shot Depth Aware Image Editing With Diffusion ModelsRishubh Parihar, Sachidanand VS, R. Venkatesh BabuICCV 2025 · 3 citations
- Camera Control for Text-to-Image Generation via Learning Viewpoint TokensXinxuan Lu, Charless Fowlkes, Alexander C. BergCVPR 2026 · 2 citations
Builds on41
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- TokenVerse: Versatile Multi-concept Personalization in Token Modulation SpaceDaniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov et al.SIGGRAPH 2025 · 12 citations
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson et al.NeurIPS 2024 · 32 citations
- XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT ModulationBowen Chen, Brynn zhao, Haomiao Sun, Li Chen et al.NeurIPS 2025 · 60 citations
- Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image GenerationAbdelrahman Eldesokey, Peter WonkaICLR 2025 · 1 citation
