Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed Reality
Taewook Ha, Woojin Cho, Dooyoung Kim, Woontack Woo
Abstract
We propose Int3DNet, a scene-aware network that predicts 3D intention areas directly from scene geometry and head-hand motion cues, enabling robust human intention prediction without explicit object-level perception. In Mixed Reality (MR), intention prediction is critical as it enables the system to anticipate user actions and respond proactively, reducing interaction delays and ensuring seamless user experiences. Our method employs a cross attention fusion of sparse motion cues and scene point clouds, offering a novel approach that directly interprets the user's spatial intention within the scene. We evaluated Int3DNet on MoGaze and CIRCLE datasets, which are public datasets for full-body human-scene interactions, showing consistent performance across time horizons of up to 1500 ms and outperforming the baselines, even in diverse and unseen scenes. Moreover, we demonstrate the usability of proposed method through a demonstration of efficient visual question answering (VQA) based on intention areas. Int3DNet provides reliable 3D intention areas derived from head-hand motion and scene geometry, thus enabling seamless interaction between humans and MR systems through proactive processing of intention areas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07753700-a2e1-4e0d-bbec-702a9b1842e0Builds on11
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion PredictionTiezheng Ma, Yongwei Nie, Chengjiang Long, Qing Zhang et al.CVPR 2022 · 150 citations
- Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine PerceptionXiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters et al.ICCV 2023 · 145 citations
- Motion Prediction using Trajectory CuesZhenguang Liu, Pengxiang Su, Shuang Wu, Xuanjing Shen et al.ICCV 2021 · 63 citations
- So Predictable! Continuous 3D Hand Trajectory Prediction in Virtual RealityNisal Menuka Gamage, Deepana Ishtaweera, Martin Weigel, Anusha WithanaUIST 2021 · 45 citations
Related papers
- Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene DescriptionAnna-Maria Halacheva, Yang Miao, Jan-Nico Zaech, Xi Wang et al.ICCV 2025 · 2 citations
- Multimodal Sense-Informed Forecasting of 3D Human MotionsZhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang et al.CVPR 2024 · 8 citations
- Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion PredictionZhenyu Lou, Qiongjie Cui, Tuo Wang, Zhenbo Song et al.NeurIPS 2024 · 10 citations
- Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue IntegrationXuejing Luo, Hee-Seung Moon, Christian Holz, Antti OulasvirtaCHI 2026 · 1 citation
- IntentMotion: Learning Intent-Aware Human Motion from Language in 3D ScenesWenfeng Song, Shi Zheng, Xinyu Zhang, Xingliang Jin et al.AAAI 2026
