Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling
Christopher Xie, Armen Avetisyan, Henry Howard-Jenkins, Yawar Siddiqui, Julian Straub, Richard A. Newcombe, Vasileios Balntas, Jakob J. Engel
摘要
We present a novel human-in-the-loop approach to estimate 3D scene layout that uses human feedback from an egocentric standpoint. We study this approach through introduction of a novel local correction task, where users identify local errors and prompt a model to automatically correct them. Building on SceneScript [3], a state-of-the-art framework for 3D scene layout estimation that leverages structured language, we propose a solution that structures this problem as "infilling", a task studied in natural language processing. We train a multi-task version of Sce-neScript that maintains performance on global predictions while significantly improving its local correction ability. We integrate this into a human-in-the-loop system, enabling a user to iteratively refine scene layout estimates via a lowfriction "one-click fix" workflow. Our system enables the final refined layout to diverge from the training distribution, allowing for more accurate modelling of complex layouts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet 等ICLR 2022 · 被引用 435 次
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 被引用 339 次
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 被引用 290 次
相关 Paper
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding 等ICLR 2026 · 被引用 74 次
- Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David AcunaCVPR 2025
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 被引用 2 次
- iPLAN: Interactive and Procedural Layout PlanningFeixiang He, Yanlong Huang, He WangCVPR 2022 · 被引用 22 次
- SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM InteractionZeyuan Huang, Cangjun Gao, Yaxian Shan, Haoxiang Hu 等UIST 2025 · 被引用 8 次
