RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-Horizon Robot Demonstration
Longxin Kou, Fei Ni, Yan Zheng, Peilong Han, Jinyi Liu, Haiqin Cui, Rui Liu, Jianye Hao
Abstract
Recent advances in robotics have produced numerous valuable large-scale demonstration datasets, yet their potential remains underutilized due to annotation limitations. Current datasets often suffer from sparse temporal annotations and inconsistent labeling granularity, particularly for complex long-horizon demonstrations. Traditional manual annotation methods are expensive and poorly scalable while existing automated methods struggle with temporal coherence and semantic richness across extended demonstrations. For this, we propose RoboAnnotatorX, a reliable annotation tool that enhances multimodal large language model to generate high-quality, context-rich annotations for complex long-horizon demonstrations. Specifically, we introduce a multi-scale token-efficient encoder
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22a6dc62-bb6b-4cb3-a157-ba589cd39ab4Cited by top-tier papers2
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang et al.CVPR 2026 · 1 citation
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong et al.CVPR 2026
Builds on13
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationYufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang et al.ICML 2024 · 227 citations
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani et al.ICML 2023 · 212 citations
- Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics TasksMurtaza Dalal, Tarun Chiruvolu, Devendra Singh Chaplot, Ruslan SalakhutdinovICLR 2024 · 86 citations
Related papers
- KISA: A Unified Keyframe Identifier and Skill Annotator for Long-Horizon Robotics DemonstrationsLongxin Kou, Fei Ni, Yan Zheng, Jinyi Liu et al.ICML 2024 · 5 citations
- SpeechCraft: A Fine-Grained Expressive Speech Dataset with Natural Language DescriptionZeyu Jin, Jia Jia, Qixin Wang, Kehan Li et al.ACM MM 2024 · 12 citations
- LoVR: A Benchmark for Long Video Retrieval in Multimodal ContextsHao Liang, Qifeng Cai, Zhaoyang Han, Hejun Dong et al.WWW 2026 · 11 citations
- FrankenMotion: Part-level Human Motion Generation and CompositionChuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger et al.CVPR 2026 · 10 citations
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
