ARNOLD: A Benchmark for Language-Grounded Task Learning With Continuous States in Realistic 3D Scenes
Ran Gong, Jiangyong Huang, Yizhou Zhao, Haoran Geng, Xiaofeng Gao, Qingyang Wu, Wensi Ai, Ziheng Zhou, Demetri Terzopoulos, Song-Chun Zhu, Baoxiong Jia, Siyuan Huang
摘要
Understanding the continuous states of objects is essential for task learning and planning in the real world. However, most existing task learning benchmarks assume discrete (e.g., binary) object goal states, which poses challenges for the learning of complex tasks and transferring learned policy from simulated environments to the real world. Furthermore, state discretization limits a robot’s ability to follow human instructions based on the grounding of actions and states. To tackle these challenges, we present ARNOLD, a benchmark that evaluates language-grounded task learning with continuous states in realistic 3D scenes. ARNOLD is comprised of 8 language-conditioned tasks that involve understanding object states and learning policies for continuous goals. To promote language-instructed learning, we provide expert demonstrations with template-generated language descriptions. We assess task performance by utilizing the latest language-conditioned policy learning models. Our results indicate that current models for language-conditioned manipulations continue to experience significant challenges in novel goal-state generalizations, scene generalizations, and object generalizations. These findings highlight the need to develop new algorithms that address this gap and underscore the potential for further research in this area. Project website: https://arnold-benchmark.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- PhyRecon: Physically Plausible Neural Scene ReconstructionJunfeng Ni, Yixin Chen, Bohan Jing, Nan Jiang 等NeurIPS 2024 · 被引用 54 次
- PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AIYandan Yang, Baoxiong Jia, Peiyuan Zhi, Siyuan HuangCVPR 2024 · 被引用 27 次
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D ScenesXiongkun Linghu, Jiangyong Huang, Ziyu Zhu, Baoxiong Jia 等ICLR 2026 · 被引用 9 次
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksKaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang 等AAAI 2026 · 被引用 6 次
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
相关 Paper
- SD2 Actor: Continuous State Decomposition Via Diffusion Embeddings for Robotic ManipulationJiayi LiICCV 2025 · 被引用 1 次
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
- Placeit3d: Language-Guided Object Placement in Real 3D ScenesAhmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi 等ICCV 2025 · 被引用 11 次
- InstructPart: Task-Oriented Part Segmentation with Instruction ReasoningZifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin 等ACL 2025 · 被引用 6 次
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationYash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang 等ICLR 2026 · 被引用 22 次
