Measuring scalar constructs in social science with LLMs
Hauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle
Abstract
Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex," but exists on a continuum between extremes. Although large language models (LLMs) are an attractive tool for measuring scalar constructs, their idiosyncratic treatment of numerical outputs raises questions of how to best apply them. We address these questions with a comprehensive evaluation of LLM-based approaches to scalar construct measurement in social science. Using multiple datasets sourced from the political science literature, we evaluate four approaches: unweighted direct pointwise scoring, aggregation of pairwise comparisons, tokenprobability-weighted pointwise scoring, and finetuning. Our study finds that pairwise comparisons made by LLMs produce better measurements than simply prompting the LLM to directly output the scores, which suffers from bunching around arbitrary numbers. However, taking the weighted mean over the token probability of scores further improves the measurements over the two previous approaches. Finally, finetuning smaller models with as few as 1,000 training pairs can match or exceed the performance of prompted LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d70f7c97-1679-4a42-a0dd-2e54c4318ea1Cited by top-tier papers4
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou et al.ACL 2026 · 51 citations
- Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsGeorg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus StrohmaierACL 2026 · 4 citations
- Counterfactual LLM-based Framework for Measuring Rhetorical StyleJingyi Qiu, Hong Chen, Zongyi LiICLR 2026 · 1 citation
- How Persuasive Is Your Context?Tu Nguyen, Kevin Du, Alexander Miserlis Hoyle, Ryan CotterellEMNLP 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
Related papers
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 31 citations
- Multilingual estimation of political-party positioning: From label aggregation to long-input TransformersDmitry Nikolaev, Tanise Ceron, Sebastian PadóEMNLP 2023 · 2 citations
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 33 citations
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
- Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and CompetenceMutaz Ayesh, Saif M. Mohammad, Nedjma OusidhoumACL 2026
