Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the "What" and Focus on the "How"!
Chen Chen, Dylan Walker, Venkatesh Saligrama
Abstract
We propose a novel supervised learning approach for political ideology prediction (PIP) that is capable of predicting out-of-distribution inputs. This problem is motivated by the fact that manual data-labeling is expensive, while self-reported labels are often scarce and exhibit significant selection bias. We propose a novel statistical model that decomposes the document embeddings into a linear superposition of two vectors; a latent neutral context vector independent of ideology, and a latent position vector aligned with ideology. We train an end-to-end model that has intermediate contextual and positional vectors as outputs. At deployment time, our model predicts labels for input documents by exclusively leveraging the predicted positional vectors. On two benchmark datasets we show that our model is capable of outputting predictions even when trained with as little as 5% biased data, and is significantly more accurate than the state-of-the-art. Through crowd-sourcing we validate the neutrality of contextual vectors, and show that context filtering results in ideological concentration, allowing for prediction on out-of-distribution examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52fca7e0-74e9-4510-8aae-3a1b07e9711cBuilds on1
Related papers
- Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and AnalysisChangyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu WangEMNLP 2022 · 4 citations
- PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderYiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, Jun YuACL 2025 · 2 citations
- Unsupervised Detection of Contextualized Embedding Bias with Application to IdeologyValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeICML 2022 · 1 citation
- An Embedding Model for Estimating Legislative Preferences from the Frequency and Sentiment of TweetsGregory Spell, Brian Guay, Sunshine Hillygus, Lawrence CarinEMNLP 2020 · 6 citations
- Unsupervised Belief Representation Learning with Information-Theoretic Variational Graph Auto-EncodersJinning Li, Huajie Shao, Dachun Sun, Ruijie Wang et al.SIGIR 2022 · 36 citations
