HumanLLM: Towards Personalized Understanding and Simulation of Human Nature
Yuxuan Lei, Tianfu Wang, Jianxun Lian, Zhengyu Hu, Defu Lian, Xing Xie
Abstract
Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior—a capability with profound implications for transforming social science research and customer-centric business insights. However, LLMs often lack a nuanced understanding of human cognition and behavior, limiting their effectiveness in social simulation and personalized applications. We posit that this limitation stems from a fundamental misalignment: standard LLM pretraining on vast, uncontextualized web data does not capture the continuous, situated context of an individual's decisions, thoughts, and behaviors over time. To bridge this gap, we introduce HumanLLM, a foundation model designed for personalized understanding and simulation of individuals. We first construct the Cognitive Genome Dataset, a large-scale corpus curated from real-world user data on platforms like Reddit, Twitter, Blogger, and Amazon. Through a rigorous, multi-stage pipeline involving data filtering, synthesis, and quality control, we automatically extract over 5.5 million user logs to distill rich profiles, behaviors, and thinking patterns. We then formulate diverse learning tasks and perform supervised fine-tuning to empower the model to predict a wide range of individualized human behaviors, thoughts, and experiences. Comprehensive evaluations demonstrate that HumanLLM achieves superior performance in predicting user actions and inner thoughts, more accurately mimics user writing styles and preferences, and generates more authentic user profiles compared to base models. Furthermore, HumanLLM shows significant gains on out-of-domain social intelligence benchmarks, indicating enhanced generalization. This work paves the way for more human-centric AI systems by advancing research in social simulation, developing personalized companions, enabling marketing intelligence through simulated customer feedback, and powering more realistic user simulation for recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Evaluating and Inducing Personality in Pre-trained Language ModelsGuangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han et al.NeurIPS 2023 · 192 citations
- Can Large Language Models Be Good Companions?: An LLM-Based Eyewear System with Conversational Common GroundZhenyu Xu, Hailin Xu, Zhouyang Lu, Yingying Zhao et al.UbiComp 2024 · 20 citations
- SocialEval: Evaluating Social Intelligence of Large Language ModelsJinfeng Zhou, Yuxuan Chen, Yihan Shi, Xuanming Zhang et al.ACL 2025 · 9 citations
Related papers
- HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive PatternsXintao Wang, Jian Yang, Weiyuan Li, Rui Xie et al.ACL 2026
- HumanLM: Simulating Users with State Alignment Beats Response ImitationShirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang et al.ICML 2026
- Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior UnderstandingKeane Ong, Wei Dai, Carol Li, Dewei Feng et al.ICLR 2026 · 8 citations
- Turning large language models into cognitive modelsMarcel Binz, Eric SchulzICLR 2024 · 99 citations
- BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded DataWenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou et al.ACL 2025
