GloCTM: Cross-Lingual Topic Modeling via a Global Context Space
Nguyen Tien Phat, Ngo Vu Minh, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen
摘要
Cross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages—a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often fail to capture deep cross-lingual semantics, resulting in loosely connected topic spaces. Moreover, these approaches often overlook the rich semantic signals embedded in multilingual pretrained representations, further limiting their ability to capture fine-grained alignment. We introduce GloCTM (Global Context Space for Cross-Lingual Topic Model), a novel framework that enforces cross-lingual topic alignment through a unified semantic space spanning the entire model pipeline. GloCTM constructs enriched input representations by expanding bag-of-words with cross-lingual lexical neighborhoods, and infers topic proportions using both local and global encoders, with their latent representations aligned through internal regularization. At the output level, the global topic-word distribution, defined over the combined vocabulary, structurally synchronizes topic meanings across languages. To further ground topics in deep semantic space, GloCTM incorporates a Centered Kernel Alignment (CKA) loss that aligns the latent topic space with multilingual contextual embeddings. Experiments across multiple benchmarks demonstrate that GloCTM significantly improves topic coherence and cross-lingual alignment, outperforming strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language ModelsMinh Chu Xuan, Tien-Phat Nguyen, Linh Ngo Van, Dinh Viet Sang 等ACL 2026 · 被引用 1 次
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen 等ICML 2026
它引用的顶会 Paper3
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
- What Makes Transfer Learning Work for Medical Images: Feature Reuse & Other FactorsChristos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg 等CVPR 2022 · 被引用 93 次
- InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingXiaobao Wu, Xinshuai Dong, Thong Nguyen, Chaoqun Liu 等AAAI 2023 · 被引用 35 次
相关 Paper
- VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu 等ACL 2021
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingYan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng 等ACL 2023 · 被引用 18 次
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang 等ACL 2021
- UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingMingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng 等CVPR 2021
- CEMTM: Contextual Embedding-based Multimodal Topic ModelingAmirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty 等EMNLP 2025
