CARD: Cross-modal Agent Framework for Generative and Editable Residential Design
Pengyu Zeng, Jun Yin, Miao Zhang, Yuqin Dai, Jizhizi Li, ZhanXiang Jin, Shuai Lu
Abstract
In recent years, architectural design automation has made significant progress, but the complexity of open-world environments continues to make residential design a challenging task, often requiring experienced architects to perform multiple iterations and human-computer interactions. Therefore, assisting ordinary users in navigating these complex environments to generate and edit residential design is crucial. In this paper, we present the CARD framework, which leverages a system of specialized crossmodal agents to adapt to complex open-world environments. The framework includes a pointbased cross-modal information representation (CMI-P) that encodes the geometry and spatial relationships of residential rooms, a crossmodal residential generation model, supported by our customized Text2FloorEdit model, that acts as the lead designer to create standardized floor plans, and an embedded expert knowledge base for evaluating whether the designs meet user requirements and residential codes, providing feedback accordingly. Finally, a 3D rendering module assists users in visualizing and understanding the layout. CARD enables cross-modal residential generation from freetext input, empowering users to adapt to complex environments without requiring specialized expertise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aac138f0-d864-4161-b239-07c812ae856dCited by top-tier papers3
- FloorPlan-LLaMa: Aligning Architects' Feedback and Domain Knowledge in Architectural Floor Plan GenerationJun Yin, Pengyu Zeng, Haoyuan Sun, Yuqin Dai et al.ACL 2025 · 11 citations
- ArchiSet: Benchmarking Editable and Consistent Single-View 3D Reconstruction of Buildings with Specific Window-to-Wall RatiosJun Yin, Pengyu Zeng, Licheng Shen, Miao Zhang et al.ICCV 2025 · 2 citations
- Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor PlansSizhong Qin, Ramon Elias Weber, Xinzheng LuCVPR 2026 · 1 citation
Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
Related papers
- Unified Vector Floorplan Generation via Markup RepresentationKaede Shiohara, Toshihiko YamasakiCVPR 2026
- Co-Layout: LLM-driven Co-optimization for Interior LayoutChucheng Xiang, Ruchao Bao, Biyin Feng, Wenzheng Wu et al.AAAI 2026
- CARD: Towards Conditional Design of Multi-agent Topological StructuresTongtong Wu, Yanming Li, Ziye Tang, Chen Jiang et al.ICLR 2026 · 7 citations
- Learning to Be a Doctor: Searching for Effective Medical Agent ArchitecturesYangyang Zhuang, Wenjia Jiang, Jiayu Zhang, Ze Yang et al.ACM MM 2025 · 1 citation
- MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasksLirong Che, Shuo Wen, Shan Huang, Chuang Wang et al.CVPR 2026 · 8 citations
