Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
Junlong Chen, Jens Grubert, Per Ola Kristensson
Abstract
As more applications of large language models (LLMs) for 3D content in immersive environments emerge, it is crucial to study user behavior to identify interaction patterns and potential barriers to guide the future design of immersive content creation and editing systems which involve LLMs. In an empirical user study with 12 participants, we combine quantitative usage data with post-experience questionnaire feedback to reveal common interaction patterns and key barriers in LLM-assisted 3D scene editing systems. We identify opportunities for improving natural language interfaces in 3D design tools and propose design recommendations. Through an empirical study, we demonstrate that LLM-assisted interactive systems can be used productively in immersive environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language ModelsPingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo et al.NeurIPS 2025 · 24 citations
- RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision PeopleXinyun Cao, Kexin Phyllis Ju, Chenglin Li, Venkatesh Potluri et al.CHI 2026 · 1 citation
- Vinedresser3D: Towards Agentic Text-guided 3D EditingYankuan Chi, Xiang Li, Zixuan Huang, James M.CVPR 2026
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey et al.CHI 2024 · 124 citations
Related papers
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR ExperienceZhizhuo Yin, Yuyang Wang, Theodoros Papatheodorou, Pan HuiIEEE VR 2024 · 25 citations
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li et al.CHI 2024 · 143 citations
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- Seeking Inspiration through Human-LLM InteractionXinrui Lin, Heyan Huang, Kaihuang Huang, Xin Shu et al.CHI 2025 · 15 citations
