Limited Linguistic Diversity in Embodied AI Datasets
Selma Liliane Wanna, Agnes Luhtaru, Jonathan Salfity, Ryan Barron, Juston Moore, Cynthia Matuszek, Mitch Pryor
Abstract
Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate these systems remain poorly documented. In this work, we present a systematic dataset audit of several widely used VLA corpora, aiming to characterize what kinds of instructions these datasets actually contain and how much linguistic variety they provide. We quantify instruction language along complementary dimensions--including lexical variety, duplication and overlap, semantic similarity, and syntactic complexity. Our analysis shows that many datasets rely on highly repetitive, template-like commands with limited structural variation, yielding a narrow distribution of instruction forms. We position these findings as descriptive documentation of the language signal available in current VLA training and evaluation data, intended to support more detailed dataset reporting, more principled dataset selection, and targeted curation or augmentation strategies that broaden language coverage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eecb51ad-e08a-422c-b130-fadb3197bbffBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
Related papers
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
- VLANeXt: Recipes for Building Strong VLA ModelsXiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang et al.ICML 2026 · 10 citations
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
- Visual Compositional TuningXindi Wu, Hee Seung Hwang, Polina Kirichenko, Esin Tureci et al.ICLR 2026 · 3 citations
- ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction DataYufan Shen, Chuwei Luo, Zhaoqing Zhu, Yang Chen et al.AAAI 2025 · 6 citations
