Enhancing Text Classification via Discovering Additional Semantic Clues from Logograms
Chen Qian, Fuli Feng, Lijie Wen, Li Lin, Tat-Seng Chua
Abstract
Text classification in low-resource languages (eg Thai) is of great practical value for some information retrieval applications (eg sentiment-analysis-based restaurant recommendation). Due to lacking large-scale corpus for learning comprehensive text representation, bilingual text classification which borrows the linguistics knowledge from a rich-resource language becomes a promising solution. Despite the success of bilingual methods, they largely ignore another source of semantic information---the writing system. Noting that most low-resource languages are phonographic languages, we argue that a logographic language (eg Chinese) can provide helpful information for improving some phonographic languages' text classification, since a logographic character (ie logogram) could represent a sememe or a whole concept, not only a phoneme or a sound. In this paper, by using a phonographic labeled corpus and its machine-translated logographic corpus both, we devise a framework to explore the central theme of utilizing logograms as a "semantic detection assistant''. Specifically, from a logographic labeled corpus, we first devise a statistical-significance-based module to pick out informative text pieces. To represent them and further reduce the effects of translation errors, our approach is equipped with Gaussian embedding whose covariances serve as reliable signals of translation errors. For a test document, all seeds' Gaussian representations are used to convolute the document and produce a logographic embedding, before being fused with its phonographic embedding for final prediction. Extensive experiments validate the effectiveness of our approach and further investigations show its generalizability and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- A Training-Free Debiasing Framework with Counterfactual Reasoning for Conversational Emotion DetectionGeng Tu, Ran Jing, Bin Liang, Min Yang et al.EMNLP 2023 · 9 citations
- Counterfactual Inference for Text Classification DebiasingChen Qian, Fuli Feng, Lijie Wen, Chunping Ma et al.ACL 2021
Builds on1
Related papers
- Exploiting Cross-Lingual Subword Similarities in Low-Resource Document ClassificationMozhi Zhang, Yoshinari Fujinuma, Jordan L. Boyd-GraberAAAI 2020 · 21 citations
- Ideography Leads Us to the Field of Cognition: A Radical-Guided Associative Model for Chinese Text ClassificationHanqing Tao, Shiwei Tong, Kun Zhang, Tong Xu et al.AAAI 2021 · 15 citations
- PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling CheckLi Huang, Junjie Li, Weiwei Jiang, Zhiyu Zhang et al.ACL 2021
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin InformationZijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng et al.ACL 2021
- Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource LanguagesXu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu et al.ACL 2022
