FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
Kewei Chen, Yayu Long, Shuai Li, Mingsheng Shang
Abstract
The powerful generalization of Vision-Language-Action (VLA) models is bottlenecked by their heavy reliance on massive, redundant, and unevenly valued datasets, hindering their widespread application. Existing model-centric optimization paths, such as model compression (which often leads to performance degradation) or policy distillation (whose products are model-dependent and lack generality), fail to fundamentally address this data-level challenge. To this end, this paper introduces FT-NCFM, a fundamentally different, data-centric generative data distillation framework. Our framework employs a self-contained Fact-Tracing (FT) engine that combines causal attribution with programmatic contrastive verification to assess the intrinsic value of samples. Guided by these assessments, an adversarial NCFM process synthesizes a model-agnostic, information-dense, and reusable data asset. Experimental results on several mainstream VLA benchmarks show that models trained on just 5% of our distilled coreset achieve a success rate of 85-90% compared with training on the full dataset, while reducing training time by over 80%. Our work demonstrates that intelligent data distillation is a highly promising new path for building efficient, high-performance VLA models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 252b4e6b-ee95-4ae3-bbb5-e975e8eaac8fBuilds on8
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen et al.ICLR 2024 · 309 citations
- Zero-Shot Robotic Manipulation with Pre-Trained Image-Editing Diffusion ModelsKevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke et al.ICLR 2024 · 284 citations
- VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot ManipulationYoupeng Wen, Junfan Lin, Yi Zhu, Jianhua Han et al.NeurIPS 2024 · 63 citations
- MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot ManipulationRongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng et al.AAAI 2026 · 56 citations
- DataMIL: Selecting Data for Robot Imitation Learning with DatamodelsShivin Dass, Alaa Khaddaj, Logan Engstrom, Aleksander Madry et al.ICLR 2026 · 34 citations
Related papers
- Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data AugmentationChenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang et al.ICML 2026 · 2 citations
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- Filtering, Distillation, and Hard Negatives for Vision-Language Pre-TrainingFilip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov et al.CVPR 2023
- Unsupervised Pretraining for Fact Verification by Language Model DistillationAdrián Bazaga, Pietro Lio, Gos MicklemICLR 2024 · 5 citations
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
