Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings
Hans William Alexander Hanley, Zakir Durumeric
Abstract
Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual settings. In this work, we present a novel, scalable, interpretable, hierarchical, and multilingual approach to clustering news articles and social media data. To do this, we first train multilingual Matryoshka embeddings that can determine story similarity at varying levels of granularity based on which subset of the dimensions of the embeddings is examined. This embedding model achieves state-of-the-art performance on the SemEval 2022 Task 8 test dataset (Pearson ρ = 0.816). Once trained, we develop an efficient hierarchical clustering algorithm that leverages the hierarchical nature of Matryoshka embeddings to identify unique news stories, narratives, and themes. We conclude by illustrating how our approach can identify and cluster stories, narratives, and overarching themes within real-world news datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e26d32b-8d06-4df4-b062-ef3e29239006Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
Related papers
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 8 citations
- ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldZiyin Zhang, Zihan Liao, Hang Yu, Peng Di et al.ICML 2026
- On Fake News Detection with LLM Enhanced Semantics MiningXiaoxiao Ma, Yuchen Zhang, Kaize Ding, Jian Yang et al.EMNLP 2024 · 23 citations
- Enhancing Event-centric News Cluster Summarization via Data Sharpening and Localization InsightsLongyin Zhang, Bowei Zou, AiTi AwACL 2025 · 1 citation
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen et al.ICLR 2026 · 40 citations
