MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
Jinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu, Weipeng Zhong, Ran Yi, Yichen Jin, Zhaoyang Lyu, Feng Zheng, Lizhuang Ma, Jiangmiao Pang
Abstract
The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are limited in terms of plausibility or alignment with the tasks. In this paper, we formulate a novel task, namely task-oriented tabletop scene generation, which poses significant challenges due to the substantial gap between high-level task instructions and the tabletop scenes. To support research on such a challenging task, we introduce MesaTask-10K, a large-scale dataset comprising approximately 10,700 synthetic tabletop scenes with manually crafted layouts that ensure realistic layouts and intricate inter-object relations. To bridge the gap between tasks and scenes, we propose a Spatial Reasoning Chain that decomposes the generation process into object inference, spatial interrelation reasoning, and scene graph construction for the final 3D layout. We present MesaTask, an LLM-based framework that utilizes this reasoning chain and is further enhanced with DPO algorithms to generate physically plausible tabletop scenes that align well with given task descriptions. Exhaustive experiments demonstrate the superior performance of MesaTask compared to baselines in generating task-conforming tabletop scenes with realistic layouts. Project page is at https://mesatask.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9366ffe9-53e9-448e-a899-78417baef34bCited by top-tier papers7
- Exploring Spatial Intelligence from a Generative PerspectiveMuzhi Zhu, Shunyao Jiang, Huanyi Zheng, Zekai Luo et al.CVPR 2026 · 1 citation
- OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference OptimizationYixuan Yang, Zhen Luo, Tongsheng Ding, Junru Lu et al.NeurIPS 2025 · 1 citation
- Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene GenerationMeisheng Zhang, Shizhao Sun, Yang Zhao, Ziyuan Liu et al.ICML 2026
- STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual SystemZhen Luo, Yixuan Yang, Xudong XU, Jinkun Hao et al.ICML 2026
- PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene GenerationWeixing Chen, Zhuoqian Feng, Yang Liu, Yexin Zhang et al.ICML 2026
Builds on15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsLongwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu et al.SIGGRAPH 2024 · 148 citations
Related papers
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial ReasoningXingjian Ran, Yixuan Li, Linning Xu, Mulin Yu et al.NeurIPS 2025 · 34 citations
- BOP-ASK: Object-Interaction Reasoning for Vision-Language ModelsVineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich et al.CVPR 2026 · 6 citations
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene ComplexityHaoming Wang, Qiyao Xue, Wei GaoCVPR 2026 · 6 citations
- From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum LearningXiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang et al.ICML 2026
- MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the MetaverseZhenyu Pan, Han LiuICLR 2026 · 49 citations
