Probabilistic Reasoning with LLMs for Privacy Risk Estimation
Jonathan Zheng, Alan Ritter, Sauvik Das, Wei (Coco) Xu
Abstract
Probabilistic reasoning is a key aspect of both human and artificial intelligence that allows for handling uncertainty and ambiguity in decision-making. In this paper, we introduce a new numerical reasoning task under uncertainty for large language models, focusing on estimating the privacy risk of user-generated documents containing privacy-sensitive information. We propose BRANCH, a new LLM methodology that estimates the k-privacy value of a text-the size of the population matching the given information. BRANCH factorizes a joint probability distribution of personal information as random variables. The probability of each factor in a population is estimated separately then combined to compute the final k-value using a Bayesian network. Our experiments show that this method successfully estimates the k-value 73% of the time, a 13% increase compared to o3-mini with chain-of-thought reasoning. We also find that LLM uncertainty is a good indicator for accuracy, as high variance predictions are 37.47% less accurate on average.
Based on the individual estimated answers, k = ⌈ 100,000 × 5% × (40% + 8%) × 20% × 10% ⌉ = 48 Reddit Post Other Documents Chatbot Message Personal Disclosure Detection Model ① Self-Disclosure Identification Python Interpreter ⑤ Probability Recombination Equation Generation Model Given the Bayesian model, the privacy risk is: k = ⌈ A × B × (C.1 + C.2) × D × E ⌉ User Document: Does Townsville have the highest inflation in the entire country? Been here 20 years. I work in Tech, but $10 for eggs is ridiculous! I don't have to deal with landlords and increasing rent, thankfully.
My daycare also increased their rate. I only have 4 months of maternity leave, so I'm looking for affordable child care options in the area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0738d6fc-5a96-452f-89cd-8a21eb09463eCited by top-tier papers2
- Large-scale online deanonymization with LLMsSimon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni et al.USENIX Security 2026 · 20 citations
- Supporting Informed Self-Disclosure: Design Recommendations for Presenting AI-Estimates of Privacy Risks to UsersIsadora Krsek, Meryl Ye, Wei Xu, Alan Ritter et al.CHI 2026 · 1 citation
Builds on25
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 13 citations
- TokUR: Token-Level Uncertainty Estimation for Large Language Model ReasoningTunyu Zhang, Haizhou Shi, Yibin Wang, Hengyi Wang et al.ICLR 2026 · 19 citations
- FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in LLMsYiyuan Li, Shichao Sun, Pengfei LiuEMNLP 2024
- CER: Confidence Enhanced Reasoning in LLMsAli Razghandi, Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani BaghshahACL 2025 · 11 citations
- Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language ModelsJavier González, Aditya V. NoriNeurIPS 2024 · 13 citations
