RPM: Reasoning-Level Personalization for Black-Box Large Language Models
Jieyong Kim, Tongyoung Kim, Soojin Yoon, Jaehyung Kim, Dongha Lee
Abstract
While black-box large language models are widely deployed, they produce generic outputs that overlook individual user preferences. Current personalization methods are fundamentally limited to response-level personalization; they only match final outputs, failing to model the underlying reasoning that connects user behavior to responses. To address this, this work introduces reasoning-level personalization as a new paradigm and proposes RPM, the first systematic framework that automatically discovers user-specific reasoning structures from raw behavioral data to guide the model's personalized inference. RPM constructs a structured model of user behavior-built from response-influential features and statistical factors-to create personalized reasoning paths and retrieve beneficial examples for guiding inference through a feature-based retrieval mechanism. Extensive experiments across four diverse tasks demonstrate that RPM consistently outperforms existing response-level methods while simultaneously enhancing both personalization performance and interpretability, providing a promising direction for black-box LLM personalization. Our code is publicly available at https://github.com/jieyong99/RPM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3deb12e1-c94d-461c-adce-f7244bbc1627Cited by top-tier papers2
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form GenerationChengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao et al.ICLR 2026 · 35 citations
- IPQA: A Benchmark for Core Intent Identification in Personalized Question AnsweringJieyong Kim, Maryam Amirizaniani, Soojin Yoon, Dongha LeeSIGIR 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- LLMRG: Improving Recommendations through Large Language Model Reasoning GraphsYan Wang, Zhixuan Chu, Xin Ouyang, Simeng Wang et al.AAAI 2024 · 47 citations
- HYDRA: Model Factorization Framework for Black-Box LLM PersonalizationYuchen Zhuang, Haotian Sun, Yue Yu, Rushi Qiang et al.NeurIPS 2024 · 79 citations
- TRIPLE: Theory-Driven Integration of Planned and Habitual Behaviors for LLM-based PersonalizationTaehyung Noh, Seungwan Jin, Haein Yeo, Kyungsik HanAAAI 2026
- Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationJieyong Kim, Hyunseo Kim, Hyunjin Cho, SeongKu Kang et al.SIGIR 2025 · 13 citations
- CBP-Tuning: Efficient Local Customization for Black-box Large Language ModelsJiaxuan Zhao, Naibin Gu, Yuchen Feng, Xiyu Liu et al.EMNLP 2025
