TAROT: Targeted Data Selection via Optimal Transport
Lan Feng, Fan Nie, Yuejiang Liu, Alexandre Alahi
Abstract
We propose TAROT, a Targeted data selection framework grounded in Optimal Transport theory. Previous targeted data selection methods primarily use influence-based greedy heuristics to enhance domain-specific performance. These methods perform well on limited, unimodal data (i.e., data following a single pattern) but become less effective as target data increases in complexity. Specifically, in multimodal distributions, these heuristics fail to account for multiple inherent patterns, leading to suboptimal data selection. This work identifies two primary factors contributing to this limitation: (i) the disproportionate impact of dominant feature components in high-dimensional influence estimation, and (ii) the restrictive linear additive assumptions inherent in greedy selection strategies. To address these challenges, TAROT incorporates whitened feature distance to mitigate dominant feature bias, offering a more reliable measure of data influence. Building on this, TAROT uses whitened feature distance to quantify and minimize the optimal transport distance between the selected data and target domains. Notably, this minimization also facilitates the estimation of optimal selection ratios. We evaluate TAROT across multiple tasks, including semantic segmentation, motion prediction, and instruction tuning. Results consistently show that TAROT outperforms state-of-the-art methods, highlighting its versatility across various deep learning tasks. Code is available at: https: //github.com/vita-epfl/TAROT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Fast Data Attribution for Text-to-Image ModelsSheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Richard Zhang et al.NeurIPS 2025 · 7 citations
- GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender SystemsTiehua Mei, Hengrui Chen, Peng Yu, Jiaqing Liang et al.KDD 2025 · 3 citations
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
Related papers
- TAROT: Task-Adaptive Refinement of LLM-prior Graphs for Few-shot Tabular LearningRuxue Shi, Yili Wang, Mengnan Du, Hangting Ye et al.KDD 2026
- A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t)Nihal Nayak, Paula Rodriguez-Diaz, Neha Hulkund, Sara Beery et al.ICML 2026 · 2 citations
- TSDS: Data Selection for Task-Specific Model FinetuningZifan Liu, Amin Karbasi, Theodoros RekatsinasNeurIPS 2024 · 35 citations
- Wasserstein Selective Transfer Learning for Cross-domain Text MiningLingyun Feng, Minghui Qiu, Yaliang Li, Haitao Zheng et al.EMNLP 2021 · 5 citations
- Probability-Polarized Optimal Transport for Unsupervised Domain AdaptationYan Wang, Chuan-Xian Ren, Yi-Ming Zhai, You-Wei Luo et al.AAAI 2024 · 8 citations
