Probabilistic Hash Embeddings for Online Learning of Categorical Features
Aodong Li, Abishek Sankararaman, Balakrishnan Narayanaswamy
摘要
We study streaming data with categorical features where the vocabulary of categorical feature values is changing and can even grow unboundedly over time. Feature hashing is commonly used as a pre-processing step to map these categorical values into a feature space of fixed size before learning their embeddings. While these methods have been developed and evaluated for offline or batch settings, in this paper we consider online settings. We show that deterministic embeddings are sensitive to the arrival order of categories and suffer from forgetting in online learning, leading to performance deterioration. To mitigate this issue, we propose a probabilistic hash embedding (PHE) model that treats hash embeddings as stochastic and applies Bayesian online learning to learn incrementally from data. Based on the structure of PHE, we derive a scalable inference algorithm to learn model parameters and infer/update the posteriors of hash embeddings and other latent variables. Our algorithm (i) can handle an evolving vocabulary of categorical items, (ii) is adaptive to new items without forgetting old items, (iii) is implementable with a bounded set of parameters that does not grow with the number of distinct observed values on the stream, and (iv) is invariant to the item arrival order. Experiments in classification, sequence modeling, and recommendation systems in online learning setups demonstrate the superior performance of PHE while maintaining high memory efficiency (consumes as low as 2→4% memory of a one-hot embedding table).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation SystemsHao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, Jiyan YangKDD 2020 · 被引用 88 次
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular DataLun Du, Fei Gao, Xu Chen, Ran Jia 等KDD 2021 · 被引用 58 次
相关 Paper
- Learning to Embed Categorical Features without Embedding Tables for RecommendationWang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi 等KDD 2021 · 被引用 46 次
- Prototypical Hash Encoding for On-the-Fly Fine-Grained Category DiscoveryHaiyang Zheng, Nan Pu, Wenjing Li, Nicu Sebe 等NeurIPS 2024 · 被引用 22 次
- Self-Distillation Dual-Memory Online Hashing with Hash Centers for Streaming Data RetrievalChong-Yu Zhang, Xin Luo, Yu-Wei Zhan, Peng-Fei Zhang 等ACM MM 2023 · 被引用 9 次
- CLEAR: Contrastive-Prototype Learning with Drift Estimation for Resource Constrained Stream MiningZhuoyi Wang, Yuqiao Chen, Chen Zhao, Yu Lin 等WWW 2021 · 被引用 21 次
- Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media RetrievalDi Wang, Quan Wang, Yaqiang An, Xinbo Gao 等SIGIR 2020 · 被引用 69 次
