Interactive Refinement of Cross-Lingual Word Embeddings
Michelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater, Jordan L. Boyd-Graber
摘要
Cross-lingual word embeddings transfer knowledge between languages: models trained on high-resource languages can predict in low-resource languages. We introduce CLIME, an interactive system to quickly refine cross-lingual word embeddings for a given classification problem. First, CLIME ranks words by their salience to the downstream task. Then, users mark similarity between keywords and their nearest neighbors in the embedding space. Finally, CLIME updates the embeddings using the annotations. We evaluate CLIME on identifying health-related text in four low-resource languages: Ilocano, Sinhalese, Tigrinya, and Uyghur. Embeddings refined by CLIME capture more nuanced word semantics and have higher test accuracy than the original embeddings. CLIME often improves accuracy faster than an active learning baseline and can be easily combined with active learning to improve results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Improving Word Translation via Two-Stage Contrastive LearningYaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen 等ACL 2022 · 被引用 32 次
- How does a Neural Network's Architecture Impact its Robustness to Noisy Labels?Jingling Li, Mozhi Zhang, Keyulu Xu, John Dickerson 等NeurIPS 2021 · 被引用 25 次
- Exploiting Cross-Lingual Subword Similarities in Low-Resource Document ClassificationMozhi Zhang, Yoshinari Fujinuma, Jordan L. Boyd-GraberAAAI 2020 · 被引用 21 次
它引用的顶会 Paper3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Exploiting Cross-Lingual Subword Similarities in Low-Resource Document ClassificationMozhi Zhang, Yoshinari Fujinuma, Jordan L. Boyd-GraberAAAI 2020 · 被引用 21 次
相关 Paper
- Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchACL 2025 · 被引用 17 次
- HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff 等ICLR 2026 · 被引用 8 次
- Cross-language Sentence Selection via Data Augmentation and Rationale TrainingYanda Chen, Chris Kedzie, Suraj Nair, Petra Galuscáková 等ACL 2021
- TEMA: Token Embeddings Mapping for Enriching Low-Resource Language ModelsRodolfo Zevallos, Núria Bel, Mireia FarrúsEMNLP 2024
- Exemplar Guided Active LearningJason S. Hartford, Kevin Leyton-Brown, Hadas Raviv, Dan Padnos 等NeurIPS 2020 · 被引用 8 次
