DIWIFT: Discovering Instance-wise Influential Features for Tabular Data
Dugang Liu, Pengxiang Cheng, Hong Zhu, Xing Tang, Yanyu Chen, Xiaoting Wang, Weike Pan, Zhong Ming, Xiuqiang He
摘要
Tabular data is one of the most common data storage formats behind many real-world web applications such as retail, banking, and e-commerce. The success of these web applications largely depends on the ability of the employed machine learning model to accurately distinguish influential features from all the predetermined features in tabular data. Intuitively, in practical business scenarios, different instances should correspond to different sets of influential features, and the set of influential features of the same instance may vary in different scenarios. However, most existing methods focus on global feature selection assuming that all instances have the same set of influential features, and few methods considering instance-wise feature selection ignore the variability of influential features in different scenarios. In this paper, we first introduce a new perspective based on the influence function for instance-wise feature selection, and give some corresponding theoretical insights, the core of which is to use the influence function as an indicator to measure the importance of an instance-wise feature. We then propose a new solution for discovering instance-wise influential features in tabular data (DIWIFT), where a self-attention network is used as a feature selection model and the value of the corresponding influence function is used as an optimization objective to guide the model. Benefiting from the advantage of the influence function, i.e., its computation does not depend on a specific architecture and can also take into account the data distribution in different scenarios, our DIWIFT has better flexibility and robustness. Finally, we conduct extensive experiments on both synthetic and real-world datasets to validate the effectiveness of our DIWIFT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation LearningTalip Ucar, Ehsan Hajiramezanali, Lindsay EdwardsNeurIPS 2021 · 被引用 189 次
- DANets: Deep Abstract Networks for Tabular Data Classification and RegressionJintai Chen, Kuanlun Liao, Yao Wan, Danny Z. Chen 等AAAI 2022 · 被引用 82 次
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He 等AAAI 2020 · 被引用 61 次
- OCT-GAN: Neural ODE-based Conditional Tabular GANsJayoung Kim, Jinsung Jeon, Jaehoon Lee, Jihyeon Hyeong 等WWW 2021 · 被引用 50 次
相关 Paper
- Can a Deep Learning Model be a Sure Bet for Tabular Prediction?Jintai Chen, Jiahuan Yan, Qiyuan Chen, Danny Z. Chen 等KDD 2024 · 被引用 4 次
- InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature InterpretationJacob Yoke Hong Si, Wendy Yusi Cheng, Michael Cooper, Rahul G. KrishnanICML 2024 · 被引用 15 次
- Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning DataIshika Agarwal, Dilek Hakkani-TurNeurIPS 2025 · 被引用 2 次
- Active feature acquisition via explainability-driven rankingOsman Berke Güney, Ketan Suhaas Saichandran, Karim Elzokm, Ziming Zhang 等ICML 2025
- Retrieval & Interaction Machine for Tabular Data PredictionJiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu 等KDD 2021 · 被引用 33 次
