SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking Services
Hongcheng Guo, Yue Wang, Shaosheng Cao, Fei Zhao, Boyang Wang, Lei Li, Liang Chen, Xinze Lyu, Zhe Xu, Yao Hu, Zhoujun Li
摘要
With the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content generation, and communication dynamics. However, recent studies focus on isolated SNS tasks rather than a comprehensive evaluation. In this paper, we introduce SNS-BENCH, specially constructed for assessing the abilities of large language models from different Social Networking Services, with a wide range of SNSrelated information. SNS-BENCH encompasses 8 different tasks such as note classification, query content relevance, and highlight words generation in comments. Finally, 6,658 questions of social media text, including subjective questions, single-choice, and multiple-choice questions, are concluded in SNS-BENCH. Further, we evaluate the performance of over 25+ current diverse LLMs on our SNS-BENCH. Models with different sizes exhibit performance variations, yet adhere to the scaling law. Moreover, we hope provide more insights to revolutionize the techniques of social network services with LLMs. https: //github.com/HC-Guo/SNS-Bench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng 等CVPR 2026 · 被引用 3 次
- Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGCLinjuan Wu, Ruiqi Zhang, Xinze Lyu, Ye Guo 等ICML 2026
它引用的顶会 Paper3
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai 等ACL 2024
相关 Paper
- SoMe: A Realistic Benchmark for LLM-based Social Media AgentsDizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu 等AAAI 2026 · 被引用 1 次
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackYouquan Li, Miao Zheng, Fan Yang, Guosheng Dong 等EMNLP 2025
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam 等ICLR 2024 · 被引用 85 次
- SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?Xudong Lu, Haohao Gao, Renshou Wu, Shuai Ren 等EMNLP 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun 等ACL 2024
