TarGATE: Target-Aware Data Selection via Token-Attenuation Gates
Xiandi Luo, Shiwei Li, Haozhao Wang, Yihao Ouyang, Zhuoqi Hu, Yichen Li, Xiao Yang, Huning Liu, Ruixuan Li
Abstract
Targeted instruction tuning requires selecting pertinent samples from massive mixed candidate datasets guided by a small reference dataset reflecting the desired capability. However, efficiently identifying high-quality data amidst noise remains challenging. To address this, we propose Target-aware GATEs (TarGATE), a simple yet effective data selection framework that leverages the model's inherent data understanding ability. These gates compute a token-level Information Retention Ratio (IRR) to attenuate the output of the feed-forward network, where the instancelevel average IRR serves as a quantitative metric for data quality. To align gates' preferences with the target task, we employ a joint optimization strategy utilizing the reference dataset and a subset of candidate data, which encourages the gates to assign higher IRRs to reference-aligned data while suppressing low-quality samples. Extensive experiments across noisy and real-world scenarios demonstrate that TarGATE outperforms related baselines. Furthermore, TarGATE exhibits superior computational efficiency and strong crossmodel transferability, enabling smaller selector to effectively curate high-quality fine-tuning data for larger foundation models. The code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3705ffa6-dcc1-40b9-9d20-256a4f4ec905Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
Related papers
- Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM FinetuningHao Cheng, Rui Zhang, Ling Li, Na Di et al.ICLR 2026
- What Makes Good Instruction-Tuning Data? An In-Context Learning PerspectiveGuangzeng Han, Xiaolei HuangACL 2026 · 1 citation
- OASIS: Online Sample Selection for Continual Instruction TuningMinjae Lee, Minhyuk Seo, Tingyu Qu, Tinne Tuytelaars et al.ACL 2026
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma et al.ICLR 2026 · 4 citations
- Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessQifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu et al.ICCV 2025 · 1 citation
