PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder
Yiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, Jun Yu
Abstract
Semantic Text Embedding is a fundamental NLP task that encodes textual content into vector representations, where proximity in the embedding space reflects semantic similarity. While existing embedding models excel at capturing general meaning, they often overlook ideological nuances, limiting their effectiveness in tasks that require an understanding of political bias. To address this gap, we introduce PRISM, the first framework designed to Produce inteRpretable polItical biaS eMbeddings. PRISM operates in two key stages: (1) Controversial Topic Bias Indicator Mining, which systematically extracts fine-grained political topics and their corresponding bias indicators from weakly labeled news data, and (2) Cross-Encoder Political Bias Embedding, which assigns structured bias scores to news articles based on their alignment with these indicators. This approach ensures that embeddings are explicitly tied to bias-revealing dimensions, enhancing both interpretability and predictive power. Through extensive experiments on two large-scale datasets, we demonstrate that PRISM outperforms stateof-the-art text embedding models in political bias classification while offering highly interpretable representations that facilitate diversified retrieval and ideological analysis. The source code is available at https://github. com/dukesun99/ACL-PRISM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb6235d2-1fef-4504-a2b1-821a303b3bc1Cited by top-tier papers2
- Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real PhotosHaodong Chen, Qiang Huang, Jiaqi Zhao, Qiuping Jiang et al.ACL 2026 · 1 citation
- Very Efficient Listwise Multimodal Reranking for Long DocumentsYiqun Sun, Pengfei Wei, Lawrence HsiehICML 2026 · 1 citation
Builds on12
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
- Toward Interpretable Semantic Textual Similarity via Optimal Transport-based Contrastive Sentence LearningSeonghyeon Lee, Dongha Lee, Seongbo Jang, Hwanjo YuACL 2022 · 25 citations
Related papers
- Unsupervised Detection of Contextualized Embedding Bias with Application to IdeologyValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeICML 2022 · 1 citation
- We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesRamy Baly, Giovanni Da San Martino, James R. Glass, Preslav NakovEMNLP 2020 · 6 citations
- Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the "What" and Focus on the "How"!Chen Chen, Dylan Walker, Venkatesh SaligramaACL 2023 · 4 citations
- Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and AnalysisChangyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu WangEMNLP 2022 · 4 citations
- Weakly Supervised Learning of Nuanced Frames for Analyzing Polarization in News MediaShamik Roy, Dan GoldwasserEMNLP 2020 · 42 citations
