GloCTM: Cross-Lingual Topic Modeling via a Global Context Space
Nguyen Tien Phat, Ngo Vu Minh, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen
Abstract
Cross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages—a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often fail to capture deep cross-lingual semantics, resulting in loosely connected topic spaces. Moreover, these approaches often overlook the rich semantic signals embedded in multilingual pretrained representations, further limiting their ability to capture fine-grained alignment. We introduce GloCTM (Global Context Space for Cross-Lingual Topic Model), a novel framework that enforces cross-lingual topic alignment through a unified semantic space spanning the entire model pipeline. GloCTM constructs enriched input representations by expanding bag-of-words with cross-lingual lexical neighborhoods, and infers topic proportions using both local and global encoders, with their latent representations aligned through internal regularization. At the output level, the global topic-word distribution, defined over the combined vocabulary, structurally synchronizes topic meanings across languages. To further ground topics in deep semantic space, GloCTM incorporates a Centered Kernel Alignment (CKA) loss that aligns the latent topic space with multilingual contextual embeddings. Experiments across multiple benchmarks demonstrate that GloCTM significantly improves topic coherence and cross-lingual alignment, outperforming strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd888d91-1364-45a8-98db-fd5d894cc937Cited by top-tier papers2
- LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language ModelsMinh Chu Xuan, Tien-Phat Nguyen, Linh Ngo Van, Dinh Viet Sang et al.ACL 2026 · 1 citation
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen et al.ICML 2026
Builds on3
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- What Makes Transfer Learning Work for Medical Images: Feature Reuse & Other FactorsChristos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg et al.CVPR 2022 · 93 citations
- InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingXiaobao Wu, Xinshuai Dong, Thong Nguyen, Chaoqun Liu et al.AAAI 2023 · 35 citations
Related papers
- VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu et al.ACL 2021
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingYan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng et al.ACL 2023 · 18 citations
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang et al.ACL 2021
- UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingMingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng et al.CVPR 2021
- CEMTM: Contextual Embedding-based Multimodal Topic ModelingAmirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty et al.EMNLP 2025
