Probabilistic Hash Embeddings for Online Learning of Categorical Features
Aodong Li, Abishek Sankararaman, Balakrishnan Narayanaswamy
Abstract
We study streaming data with categorical features where the vocabulary of categorical feature values is changing and can even grow unboundedly over time. Feature hashing is commonly used as a pre-processing step to map these categorical values into a feature space of fixed size before learning their embeddings. While these methods have been developed and evaluated for offline or batch settings, in this paper we consider online settings. We show that deterministic embeddings are sensitive to the arrival order of categories and suffer from forgetting in online learning, leading to performance deterioration. To mitigate this issue, we propose a probabilistic hash embedding (PHE) model that treats hash embeddings as stochastic and applies Bayesian online learning to learn incrementally from data. Based on the structure of PHE, we derive a scalable inference algorithm to learn model parameters and infer/update the posteriors of hash embeddings and other latent variables. Our algorithm (i) can handle an evolving vocabulary of categorical items, (ii) is adaptive to new items without forgetting old items, (iii) is implementable with a bounded set of parameters that does not grow with the number of distinct observed values on the stream, and (iv) is invariant to the item arrival order. Experiments in classification, sequence modeling, and recommendation systems in online learning setups demonstrate the superior performance of PHE while maintaining high memory efficiency (consumes as low as 2→4% memory of a one-hot embedding table).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb5a3bf2-c410-4177-b62a-ec456213345cBuilds on11
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Compositional Embeddings Using Complementary Partitions for Memory-Efficient Recommendation SystemsHao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, Jiyan YangKDD 2020 · 88 citations
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular DataLun Du, Fei Gao, Xu Chen, Ran Jia et al.KDD 2021 · 58 citations
Related papers
- Learning to Embed Categorical Features without Embedding Tables for RecommendationWang-Cheng Kang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi et al.KDD 2021 · 46 citations
- Prototypical Hash Encoding for On-the-Fly Fine-Grained Category DiscoveryHaiyang Zheng, Nan Pu, Wenjing Li, Nicu Sebe et al.NeurIPS 2024 · 22 citations
- Self-Distillation Dual-Memory Online Hashing with Hash Centers for Streaming Data RetrievalChong-Yu Zhang, Xin Luo, Yu-Wei Zhan, Peng-Fei Zhang et al.ACM MM 2023 · 9 citations
- CLEAR: Contrastive-Prototype Learning with Drift Estimation for Resource Constrained Stream MiningZhuoyi Wang, Yuqiao Chen, Chen Zhao, Yu Lin et al.WWW 2021 · 21 citations
- Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media RetrievalDi Wang, Quan Wang, Yaqiang An, Xinbo Gao et al.SIGIR 2020 · 69 citations
