SYNTHIA: Novel Concept Design with Affordance Composition
Hyeonjeong Ha, Xiaomeng Jin, Jeonghwan Kim, Jiateng Liu, Zhenhailong Wang, Khanh Duy Nguyen, Ansel Blume, Nanyun Peng, Kai-Wei Chang, Heng Ji
Abstract
Text-to-image (T2I) models enable rapid concept design, making them widely used in AI-driven design. While recent studies focus on generating semantic and stylistic variations of given design concepts, functional coherencethe integration of multiple affordances into a single coherent concept-remains largely overlooked. In this paper, we introduce SYNTHIA, a framework for generating novel, functionally coherent designs based on desired affordances. Our approach leverages a hierarchical concept ontology that decomposes concepts into parts and affordances, serving as a crucial building block for functionally coherent design. We also develop a curriculum learning scheme based on our ontology that contrastively fine-tunes T2I models to progressively learn affordance composition while maintaining visual novelty. To elaborate, we (i) gradually increase affordance distance, guiding models from basic concept-affordance association to complex affordance compositions that integrate parts of distinct affordances into a single, coherent form, and (ii) enforce visual novelty by employing contrastive objectives to push learned representations away from existing concepts. Experimental results show that SYNTHIA outperforms state-of-the-art T2I models, demonstrating absolute gains of 25.1% and 14.7% for novelty and functional coherence in human evaluation, respectively. Code is available at https://github.com/HyeonjeongHa/SYNTHIA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64571499-18d5-4856-8408-03fb01a6ffdaBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion PathsZeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu et al.NeurIPS 2023 · 201 citations
- InstantBooth: Personalized Text-to-Image Generation without Test-Time FinetuningJing Shi, Wei Xiong, Zhe Lin, Hyun Joon JungCVPR 2024 · 115 citations
Related papers
- StyleT2I: Toward Compositional and High-Fidelity Text-to-Image SynthesisZhiheng Li, Martin Renqiang Min, Kai Li, Chenliang XuCVPR 2022 · 38 citations
- ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion ModelsMaitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou YangAAAI 2024 · 16 citations
- Text-to-Image Generation with Multi-modal Knowledge Graph Construction and RetrievalJiawei Meng, Zhengmao Yang, Zhiqiang Liu, Shaokai Chen et al.ACM MM 2025
- Concept-Guided Tokenization: Closing the Gap Between Reconstruction and GenerationYunqiao Yang, Haokun Lin, Guanzhong Wu, Ying WeiICML 2026
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
