Unsupervised Generative Feature Transformation via Graph Contrastive Pre-training and Multi-objective Fine-tuning
Wangyang Ying, Dongjie Wang, Xuanming Hu, Yuanchun Zhou, Charu C. Aggarwal, Yanjie Fu
摘要
Feature transformation is to derive a new feature set from original features to augment the AI power of data. In many science domains such as material performance screening, while feature transformation can model material formula interactions and compositions and discover performance drivers, supervised labels are collected from expensive and lengthy experiments. This issue motivates an Unsupervised Feature Transformation Learning (UFTL) problem. Prior literature, such as manual transformation, supervised feedback guided search, and PCA, either relies on domain knowledge or expensive supervised feedback, or suffers from large search space, or overlooks non-linear feature-feature interactions. UFTL imposes a major challenge on existing methods: how to design a new unsupervised paradigm that captures complex feature interactions and avoids large search space? To fill this gap, we connect graph, contrastive, and generative learning to develop a measurementpretrain-finetune paradigm for UFTL. For unsupervised feature set utility measurement, we propose a feature value consistency preservation perspective and develop a mean discounted cumulative gain like unsupervised metric to evaluate feature set utility. For unsupervised feature set representation pretraining, we regard a feature set as a feature-feature interaction graph, and develop an unsupervised graph contrastive learning encoder to embed feature sets into vectors. For generative transformation finetuning, we regard a feature set as a feature cross sequence and feature transformation as sequential generation. We develop a deep generative feature transformation model that coordinates the pretrained feature set encoder and the gradient information extracted from a feature set utility evaluator to optimize a transformed feature generator. Finally, we conduct extensive experiments to demonstrate the effectiveness, efficiency, traceability, and explicitness of our framework. Our code and data are available at https://shorturl.at/pKQU5 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Evolutionary Large Language Model for Automated Feature TransformationNanxu Gong, Chandan K. Reddy, Wangyang Ying, Haifeng Chen 等AAAI 2025 · 被引用 39 次
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuningHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu 等NeurIPS 2025 · 被引用 17 次
- Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature TransformationNanxu Gong, Zijun Li, Sixun Dong, Haoyue Bai 等NeurIPS 2025 · 被引用 15 次
- Efficient Post-Training Refinement of Latent Reasoning in Large Language ModelsXinyuan Wang, Dongjie Wang, Wangyang Ying, Haoyue Bai 等AAAI 2026 · 被引用 6 次
- Brownian Bridge Augmented Surrogate Simulation and Injection Planning for Geological CO2 StorageHaoyue Bai, Guodong Chen, Wangyang Ying, Xinyuan Wang 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor 等CHI 2021 · 被引用 77 次
- Reinforcement-Enhanced Autoregressive Feature Transformation: Gradient-steered Search in Continuous Space for Postfix ExpressionsDongjie Wang, Meng Xiao, Min Wu, Pengfei Wang 等NeurIPS 2023 · 被引用 34 次
- Group-wise Reinforcement Feature Generation for Optimal and Explainable Representation Space ReconstructionDongjie Wang, Yanjie Fu, Kunpeng Liu, Xiaolin Li 等KDD 2022 · 被引用 26 次
相关 Paper
- Uncovering Capabilities of Model Pruning in Graph Contrastive LearningJunran Wu, Xueyuan Chen, Shangzhe LiACM MM 2024 · 被引用 2 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed GraphsMingqian Ding, Jianjun Li, Zhiyuan Ma, Liwei Zhang 等WWW 2026
- Let Invariant Rationale Discovery Inspire Graph Contrastive LearningSihang Li, Xiang Wang, An Zhang, Yingxin Wu 等ICML 2022 · 被引用 117 次
- FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node FeaturesJitao Zhao, Di Jin, Meng Ge, Lianze Shan 等NeurIPS 2024 · 被引用 20 次
