QueST: Self-Supervised Skill Abstractions for Learning Continuous Control
Atharva Mete, Haotian Xue, Albert Wilcox, Yongxin Chen, Animesh Garg
Abstract
Generalization capabilities, or rather a lack thereof, is one of the most important unsolved problems in the field of robot learning, and while several large scale efforts have set out to tackle this problem, unsolved it remains. In this paper, we hypothesize that learning temporal action abstractions using latent variable models (LVMs), which learn to map data to a compressed latent space and back, is a promising direction towards low-level skills that can readily be used for new tasks. Although several works have attempted to show this, they have generally been limited by architectures that do not faithfully capture shareable representations. To address this we present Quantized Skill Transformer (QueST), which learns a larger and more flexible latent encoding that is more capable of modeling the breadth of low-level skills necessary for a variety of tasks. To make use of this extra flexibility, QueST imparts causal inductive bias from the action sequence data into the latent space, leading to more semantically useful and transferable representations. We compare to state-of-the-art imitation learning and LVM baselines and see that QueST's architecture leads to strong performance on several multitask and few-shot learning benchmarks. Further results and videos are available at https://quest-model.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c73440b2-7cfc-49f5-816d-b7ca37ea6de3Cited by top-tier papers27
- Real-Time Execution of Action Chunking Flow PoliciesKevin Black, Manuel Y. Galliker, Sergey LevineNeurIPS 2025 · 280 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He et al.NeurIPS 2025 · 87 citations
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in RobotsLikui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen et al.CVPR 2026 · 18 citations
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion RepresentationsShichao Fan, Kun Wu, Zhengping Che, Xinhua Wang et al.ICML 2026 · 16 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
Related papers
- PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in ControlRuijie Zheng, Ching-An Cheng, Hal Daumé III, Furong Huang et al.ICML 2024 · 17 citations
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song et al.NeurIPS 2022 · 43 citations
- STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector QuantizationHao Li, Qi Lv, Rui Shao, Xiang Deng et al.ICML 2025
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider et al.ICML 2023 · 7 citations
- LatentVLA: Taming Latent Space for Generalizable and Long-Horizon Bimanual ManipulationJunming WangAAAI 2026 · 1 citation
