Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction
Hao Phung, Hadar Averbuch-Elor
Abstract
Fig. 1. Our approach transforms rasterized floorplan images to vectorized format, reconstructing both its structure and semantics. We illustrate * results on held-out CubiCasa5K [Kalervo et al. 2019] test samples (left). The colors denote unique semantic categories (e.g., Outdoor, Bedroom, bath, and entry). Additionally, we highlight our model's generalization capabilities over complicated real-world floorplan images from WAFFLE [Ganon et al. 2025] (right). * 3D visualizations are constructed by extending the 2D boundaries vertically.
Reconstructing a structured vector-graphics representation from a rasterized floorplan image is typically an important prerequisite for computational tasks involving floorplans such as automated understanding or CAD workflows. However, existing techniques struggle in faithfully generating the structure and semantics conveyed by complex floorplans that depict large indoor spaces with many rooms and a varying numbers of polygon corners. To this end, we propose Raster2Seq, framing floorplan reconstruction as a sequence-to-sequence task in which floorplan elements-such as rooms, windows, and doors-are represented as labeled polygon sequences that jointly encode geometry and semantics. Our approach introduces an autoregressive decoder that learns to predict the next corner conditioned on image features and previously generated corners using guidance from learnable anchors. These anchors represent spatial coordinates in image space, hence allowing for effectively directing the attention mechanism to focus on informative image regions. By embracing the autoregressive mechanism, our method offers flexibility in the output format, enabling for efficiently handling complex floorplans with numerous rooms and diverse polygon structures. Our method achieves state-of-the-art performance on standard benchmarks such as Structure3D, CubiCasa5K, and Raster2Graph, while also demonstrating strong generalization to more challenging datasets like WAFFLE, which contain diverse room structures and complex geometric variations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9431bb39-18f4-4f3f-a79f-e825861a511aBuilds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided AttentionZhiliang Zeng, Xianzhi Li, Ying Kin Yu, Chi-Wing FuICCV 2019 · 125 citations
Related papers
- Connecting the Dots: Floorplan Reconstruction Using Two-Level QueriesYuanwen Yue, Theodora Kontogianni, Konrad Schindler, Francis EngelmannCVPR 2023
- VectorFloorSeg: Two-Stream Graph Attention Network for Vectorized Roughcast Floorplan SegmentationBingchen Yang, Haiyong Jiang, Hao Pan, Jun XiaoCVPR 2023
- HEAT: Holistic Edge Attention Transformer for Structured ReconstructionJiacheng Chen, Yiming Qian, Yasutaka FurukawaCVPR 2022 · 35 citations
- HouseDiffusion: Vector Floorplan Generation via a Diffusion Model with Discrete and Continuous DenoisingMohammad Amin Shabani, Sepidehsadat Hosseini, Yasutaka FurukawaCVPR 2023
- Residential Floor Plan Recognition and ReconstructionXiaolei Lv, Shengchu Zhao, Xinyang Yu, Binqiang ZhaoCVPR 2021
