PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset Selection
Suraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff A. Bilmes, Rishabh K. Iyer
摘要
With ever-increasing dataset sizes, subset selection techniques are becoming increasingly important for a plethora of tasks. It is often necessary to guide the subset selection to achieve certain desiderata, which includes focusing or targeting certain data points, while avoiding others. Examples of such problems include: i) targeted learning, where the goal is to find subsets with rare classes or rare attributes on which the model is underperforming, and ii) guided summarization, where data (e.g., image collection, text, document or video) is summarized for quicker human consumption with specific additional user intent. Motivated by such applications, we present PRISM, a rich class of PaRameterIzed Submodular information Measures. Through novel functions and their parameterizations, PRISM offers a variety of modeling capabilities that enable a trade-off between desired qualities of a subset like diversity or representation and similarity/dissimilarity with a set of data points. We demonstrate how PRISM can be applied to the two real-world problems mentioned above, which require guided subset selection. In doing so, we show that PRISM interestingly generalizes some past work, therein reinforcing its broad utility. Through extensive experiments on diverse datasets, we demonstrate the superiority of PRISM over the state-of-the-art in targeted learning and in guided imagecollection summarization. PRISM is available as a part of the SUBMODLIB ( https://github.com/decile-team/submodlib ) and TRUST ( https://github.com/decile-team/trust ) toolkits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Contributing Dimension Structure of Deep Feature for Coreset SelectionZhijing Wan, Zhixiang Wang, Yuran Wang, Zheng Wang 等AAAI 2024 · 被引用 11 次
- SCoRe: Submodular Combinatorial Representation LearningAnay Majee, Suraj Kothawade, Krishnateja Killamsetty, Rishabh K. IyerICML 2024 · 被引用 7 次
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 被引用 5 次
- DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent AdaptationSuraj Kothawade, Anmol Reddy Mekala, D. Chandra Sekhara Hetha Havya, Mayank Kothyari 等ACL 2023 · 被引用 4 次
- Evolution-aware VAriance (EVA) Coreset Selection for Medical Image ClassificationYuxin Hong, Xiao Zhang, Xin Zhang, Joey Tianyi ZhouACM MM 2024 · 被引用 4 次
它引用的顶会 Paper6
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 被引用 300 次
- Implicit Diversity in Image SummarizationL. Elisa Celis, Vijay KeswaniCSCW 2020 · 被引用 26 次
- The Online Submodular Cover ProblemAnupam Gupta, Roie LevinSODA 2020 · 被引用 20 次
相关 Paper
- SIMILAR: Submodular Information Measures Based Active Learning In Realistic ScenariosSuraj Kothawade, Nathan Beck, KrishnaTeja Killamsetty, Rishabh K. IyerNeurIPS 2021 · 被引用 138 次
- Measures of diversity and space-filling designs for categorical dataCédric Malherbe, Emilio Domínguez-Sánchez, Merwan Barlier, Igor Colin 等ICML 2024
- Fairness in Streaming Submodular Maximization: Algorithms and HardnessMarwa El Halabi, Slobodan Mitrovic, Ashkan Norouzi-Fard, Jakab Tardos 等NeurIPS 2020 · 被引用 65 次
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 被引用 22 次
- Dynamic Submodular MaximizationMorteza MonemizadehNeurIPS 2020 · 被引用 13 次
