Unsupervised Generative Feature Transformation via Graph Contrastive Pre-training and Multi-objective Fine-tuning
Wangyang Ying, Dongjie Wang, Xuanming Hu, Yuanchun Zhou, Charu C. Aggarwal, Yanjie Fu
Abstract
Feature transformation is to derive a new feature set from original features to augment the AI power of data. In many science domains such as material performance screening, while feature transformation can model material formula interactions and compositions and discover performance drivers, supervised labels are collected from expensive and lengthy experiments. This issue motivates an Unsupervised Feature Transformation Learning (UFTL) problem. Prior literature, such as manual transformation, supervised feedback guided search, and PCA, either relies on domain knowledge or expensive supervised feedback, or suffers from large search space, or overlooks non-linear feature-feature interactions. UFTL imposes a major challenge on existing methods: how to design a new unsupervised paradigm that captures complex feature interactions and avoids large search space? To fill this gap, we connect graph, contrastive, and generative learning to develop a measurementpretrain-finetune paradigm for UFTL. For unsupervised feature set utility measurement, we propose a feature value consistency preservation perspective and develop a mean discounted cumulative gain like unsupervised metric to evaluate feature set utility. For unsupervised feature set representation pretraining, we regard a feature set as a feature-feature interaction graph, and develop an unsupervised graph contrastive learning encoder to embed feature sets into vectors. For generative transformation finetuning, we regard a feature set as a feature cross sequence and feature transformation as sequential generation. We develop a deep generative feature transformation model that coordinates the pretrained feature set encoder and the gradient information extracted from a feature set utility evaluator to optimize a transformed feature generator. Finally, we conduct extensive experiments to demonstrate the effectiveness, efficiency, traceability, and explicitness of our framework. Our code and data are available at https://shorturl.at/pKQU5 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25d08ac3-6742-4178-a41e-5be326950d2bCited by top-tier papers9
- Evolutionary Large Language Model for Automated Feature TransformationNanxu Gong, Chandan K. Reddy, Wangyang Ying, Haifeng Chen et al.AAAI 2025 · 39 citations
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuningHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu et al.NeurIPS 2025 · 17 citations
- Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature TransformationNanxu Gong, Zijun Li, Sixun Dong, Haoyue Bai et al.NeurIPS 2025 · 15 citations
- Efficient Post-Training Refinement of Latent Reasoning in Large Language ModelsXinyuan Wang, Dongjie Wang, Wangyang Ying, Haoyue Bai et al.AAAI 2026 · 6 citations
- Brownian Bridge Augmented Surrogate Simulation and Injection Planning for Geological CO2 StorageHaoyue Bai, Guodong Chen, Wangyang Ying, Xinyuan Wang et al.AAAI 2026 · 5 citations
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor et al.CHI 2021 · 77 citations
- Reinforcement-Enhanced Autoregressive Feature Transformation: Gradient-steered Search in Continuous Space for Postfix ExpressionsDongjie Wang, Meng Xiao, Min Wu, Pengfei Wang et al.NeurIPS 2023 · 34 citations
- Group-wise Reinforcement Feature Generation for Optimal and Explainable Representation Space ReconstructionDongjie Wang, Yanjie Fu, Kunpeng Liu, Xiaolin Li et al.KDD 2022 · 26 citations
Related papers
- Uncovering Capabilities of Model Pruning in Graph Contrastive LearningJunran Wu, Xueyuan Chen, Shangzhe LiACM MM 2024 · 2 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed GraphsMingqian Ding, Jianjun Li, Zhiyuan Ma, Liwei Zhang et al.WWW 2026
- Let Invariant Rationale Discovery Inspire Graph Contrastive LearningSihang Li, Xiang Wang, An Zhang, Yingxin Wu et al.ICML 2022 · 117 citations
- FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node FeaturesJitao Zhao, Di Jin, Meng Ge, Lianze Shan et al.NeurIPS 2024 · 20 citations
