DIWIFT: Discovering Instance-wise Influential Features for Tabular Data
Dugang Liu, Pengxiang Cheng, Hong Zhu, Xing Tang, Yanyu Chen, Xiaoting Wang, Weike Pan, Zhong Ming, Xiuqiang He
Abstract
Tabular data is one of the most common data storage formats behind many real-world web applications such as retail, banking, and e-commerce. The success of these web applications largely depends on the ability of the employed machine learning model to accurately distinguish influential features from all the predetermined features in tabular data. Intuitively, in practical business scenarios, different instances should correspond to different sets of influential features, and the set of influential features of the same instance may vary in different scenarios. However, most existing methods focus on global feature selection assuming that all instances have the same set of influential features, and few methods considering instance-wise feature selection ignore the variability of influential features in different scenarios. In this paper, we first introduce a new perspective based on the influence function for instance-wise feature selection, and give some corresponding theoretical insights, the core of which is to use the influence function as an indicator to measure the importance of an instance-wise feature. We then propose a new solution for discovering instance-wise influential features in tabular data (DIWIFT), where a self-attention network is used as a feature selection model and the value of the corresponding influence function is used as an optimization objective to guide the model. Benefiting from the advantage of the influence function, i.e., its computation does not depend on a specific architecture and can also take into account the data distribution in different scenarios, our DIWIFT has better flexibility and robustness. Finally, we conduct extensive experiments on both synthetic and real-world datasets to validate the effectiveness of our DIWIFT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation LearningTalip Ucar, Ehsan Hajiramezanali, Lindsay EdwardsNeurIPS 2021 · 189 citations
- DANets: Deep Abstract Networks for Tabular Data Classification and RegressionJintai Chen, Kuanlun Liao, Yao Wan, Danny Z. Chen et al.AAAI 2022 · 82 citations
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He et al.AAAI 2020 · 61 citations
- OCT-GAN: Neural ODE-based Conditional Tabular GANsJayoung Kim, Jinsung Jeon, Jaehoon Lee, Jihyeon Hyeong et al.WWW 2021 · 50 citations
Related papers
- Can a Deep Learning Model be a Sure Bet for Tabular Prediction?Jintai Chen, Jiahuan Yan, Qiyuan Chen, Danny Z. Chen et al.KDD 2024 · 4 citations
- InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature InterpretationJacob Yoke Hong Si, Wendy Yusi Cheng, Michael Cooper, Rahul G. KrishnanICML 2024 · 15 citations
- Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning DataIshika Agarwal, Dilek Hakkani-TurNeurIPS 2025 · 2 citations
- Active feature acquisition via explainability-driven rankingOsman Berke Güney, Ketan Suhaas Saichandran, Karim Elzokm, Ziming Zhang et al.ICML 2025
- Retrieval & Interaction Machine for Tabular Data PredictionJiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu et al.KDD 2021 · 33 citations
