Discrete Latent Variable Representations for Low-Resource Text Classification
Shuning Jin, Sam Wiseman, Karl Stratos, Karen Livescu
摘要
While much work on deep latent variable models of text uses continuous latent variables, discrete latent variables are interesting because they are more interpretable and typically more space efficient. We consider several approaches to learning discrete latent variable models for text in the case where exact marginalization over these variables is intractable. We compare the performance of the learned representations as features for lowresource document and sentence classification. Our best models outperform the previous best reported results with continuous representations in these low-resource settings, while learning significantly more compressed representations. Interestingly, we find that an amortized variant of Hard EM performs particularly well in the lowest-resource regimes. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang 等NeurIPS 2023 · 被引用 151 次
- Latent Diffusion Energy-Based Model for Interpretable Text ModellingPeiyu Yu, Sirui Xie, Xiaojian Ma, Baoxiong Jia 等ICML 2022 · 被引用 105 次
- Latent Space Energy-Based Model of Symbol-Vector Coupling for Text Generation and ClassificationBo Pang, Ying Nian WuICML 2021 · 被引用 19 次
- StreamHover: Livestream Transcript Summarization and AnnotationSangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui 等EMNLP 2021 · 被引用 18 次
- I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?Yuhang Liu, Dong Gong, Yichao Cai, Erdun Gao 等ICLR 2026 · 被引用 17 次
相关 Paper
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual InformationKarl Stratos, Sam WisemanICML 2020 · 被引用 9 次
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
- Learning Semantic Textual Similarity via Topic-informed Discrete Latent VariablesErxin Yu, Lan Du, Yuan Jin, Zhepei Wei 等EMNLP 2022 · 被引用 5 次
- latent-GLAT: Glancing at Latent Variables for Parallel Text GenerationYu Bao, Hao Zhou, Shujian Huang, Dongqi Wang 等ACL 2022
- Anchor & Transform: Learning Sparse Embeddings for Large VocabulariesPaul Pu Liang, Manzil Zaheer, Yuan Wang, Amr AhmedICLR 2021 · 被引用 15 次
