Language Concept Erasure for Language-invariant Dense Retrieval
Zhiqi Huang, Puxuan Yu, Shauli Ravfogel, James Allan
Abstract
Multilingual models aim for language-invariant representations but still prominently encode language identity. This, along with the scarcity of high-quality parallel retrieval data, limits their performance in retrieval. We introduce LANCER, a multi-task learning framework that improves language-invariant dense retrieval by reducing language-specific signals in the embedding space. Leveraging the notion of linear concept erasure, we design a loss function that penalizes cross-correlation between representations and their language labels. LANCER leverages only English retrieval data and general multilingual corpora, training models to focus on language-invariant retrieval by semantic similarity without necessitating a vast parallel corpus. Experimental results on various datasets show our method consistently improves over baselines, with extensive analyses demonstrating greater language agnosticism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ef11084-a758-4a93-a08d-7e81b302ac1eCited by top-tier papers3
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 4 citations
- LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity RemovalDongjun Kim, Jeongho Yoon, Chanjun Park, Heuiseok LimACL 2026
- The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept ErasureYu Fan, Yang Tian, Shauli Ravfogel, Mrinmaya Sachan et al.EMNLP 2025
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsZhihui Xie, Handong Zhao, Tong Yu, Shuai LiEMNLP 2022 · 3 citations
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan et al.EMNLP 2025 · 2 citations
- Language-Agnostic Visual-Semantic EmbeddingsJonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, Rodrigo C. BarrosICCV 2019 · 56 citations
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua et al.EMNLP 2020 · 39 citations
- Improving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentSeongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon et al.ICLR 2026 · 4 citations
