MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
Yining Hong, Zishuo Zheng, Peihao Chen, Yian Wang, Junyan Li, Chuang Gan
2024年份
16顶会引用
摘要
You are an AI assistant / task generator in the room. You need to generate a task in the scene. Demonstration: For Room 1: [Few shot example] Generate similar responses for Room 2. Response : For Room 2: Q: Is the donut ready to eat? t1 input: Q + I see a donut. output: <select> [Choose donut] t2 input: Q + I see a donut. <select> output: <touch> [tactile] [temperature] t3 input: Q + I see a donut. <select> <touch> [tactile] [temperature] output: It is hard, cold and not ready to eat.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 等CVPR 2026 · 被引用 171 次
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingYunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert 等NeurIPS 2024 · 被引用 56 次
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao 等NeurIPS 2025 · 被引用 30 次
- LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and RenderingDelin Qu, Qizhi Chen, Pingrui Zhang, Xianqiang Gao 等NeurIPS 2024 · 被引用 6 次
- HIS-GPT: Towards 3D Human-In-Scene Multimodal UnderstandingJiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
相关 Paper
- SEED-Bench: Benchmarking Multimodal Large Language ModelsBohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang 等CVPR 2024
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange 等AAAI 2022 · 被引用 251 次
- HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionYuan Wang, Yali Li, Xiang Li, Shengjin WangCVPR 2025
- Language-driven Grasp DetectionVuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen 等CVPR 2024
- Visual Room RearrangementLuca Weihs, Matt Deitke, Aniruddha Kembhavi, Roozbeh MottaghiCVPR 2021
