RepMatch: Quantifying Cross-Instance Similarities in Representation Space
Mohammad Modarres, Sina Abbasi, Mohammad Taher Pilehvar
摘要
Advances in dataset analysis techniques have enabled more sophisticated approaches to analyzing and characterizing training data instances, often categorizing data based on attributes such as "difficulty". In this work, we introduce RepMatch, a novel method that characterizes data through the lens of similarity. RepMatch quantifies the similarity between subsets of training instances by comparing the knowledge encoded in models trained on them, overcoming the limitations of existing analysis methods that focus solely on individual instances and are restricted to within-dataset analysis. Our framework allows for a broader evaluation, enabling similarity comparisons across arbitrary subsets of instances, supporting both dataset-to-dataset and instance-to-dataset analyses. We validate the effectiveness of Rep-Match across multiple NLP tasks, datasets, and models. Through extensive experimentation, we demonstrate that RepMatch can effectively compare datasets, identify more representative subsets of a dataset (that lead to better performance than randomly selected subsets of equivalent size), and uncover heuristics underlying the construction of some challenge datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 被引用 398 次
相关 Paper
- Narrowing the Gap between Supervised and Unsupervised Sentence Representation Learning with Large Language ModelMingxin Li, Richong Zhang, Zhijie Nie, Yongyi MaoAAAI 2024 · 被引用 1 次
- ReSi: A Comprehensive Benchmark for Representational Similarity MeasuresMax Klabunde, Tassilo Wald, Tobias Schumacher, Klaus H. Maier-Hein 等ICLR 2025
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 被引用 8 次
- Comparing Text Representations: A Theory-Driven ApproachGregory Yauney, David MimnoEMNLP 2021
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 被引用 105 次
