StringLLM: Understanding the String Processing Capability of Large Language Models
Xilong Wang, Hao Fu, Jindong Wang, Neil Zhenqiang Gong
Abstract
String processing, which mainly involves the analysis and manipulation of strings, is a fundamental component of modern computing. Despite the significant advancements of large language models (LLMs) in various natural language processing (NLP) tasks, their capability in string processing remains underexplored and underdeveloped. To bridge this gap, we present a comprehensive study of LLMs' string processing capability. In particular, we first propose StringLLM, a method to construct datasets for benchmarking string processing capability of LLMs. We use StringLLM to build a series of datasets, referred to as StringBench. It encompasses a wide range of string processing tasks, allowing us to systematically evaluate LLMs' performance in this area. Our evaluations indicate that LLMs struggle with accurately processing strings compared to humans. To uncover the underlying reasons for this limitation, we conduct an in-depth analysis and subsequently propose an effective approach that significantly enhances LLMs' string processing capability via fine-tuning. This work provides a foundation for future research to understand LLMs' string processing capability. Our code and data are available at https://github.com/wxl-lxw/StringLLM.
Work done when Xilong Wang was an intern at Li Auto mentored by Hao Fu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9516197-e06e-4a6f-855c-2662fc7a85ccCited by top-tier papers1
Ask how each one uses itBuilds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- OpenCoder: The Open Cookbook for Top-Tier Code Large Language ModelsSiming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu et al.ACL 2025
Related papers
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu et al.ACL 2024
- Large Language Models Are No Longer Shallow ParsersYuanhe Tian, Fei Xia, Yan SongACL 2024
- Unnatural Languages Are Not Bugs but Features for LLMsKeyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni et al.ICML 2025
- TraceLLM: Evaluating and Exploring Large Language Models on Trace Analysis in Microservice-based Web ApplicationsTong Zhou, Xin Peng, Jie Zhang, Chaofeng Sha et al.WWW 2026
- Large Language Models for Code Analysis: Do LLMs Really Do Their Job?Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu et al.USENIX Security 2024 · 110 citations
