TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
Honghua Dong, Jiacheng Yang, Xun Deng, Yuhe Jiang, Gennady Pekhimenko, Fan Long, Xujie Si
Abstract
Type inference for dynamic languages like Python is a persistent challenge in software engineering. While large language models (LLMs) have shown promise in code understanding, their type inference capabilities remain underexplored. We introduce TYPYBENCH, a benchmark designed to evaluate LLMs' type inference across entire Python repositories. TYPYBENCH features two novel metrics: TYPESIM, which captures nuanced semantic relationships between predicted and ground truth types, and TYPECHECK, which assesses type consistency across codebases. Our evaluation of various LLMs on a curated dataset of 50 high-quality Python repositories reveals that, although LLMs achieve decent TYPESIM scores, they struggle with complex nested types and exhibit significant type consistency errors. These findings suggest that future research should shift focus from improving type similarity to addressing repository-level consistency. TYPY-BENCH provides a foundation for this new direction, offering insights into model performance across different type complexities and usage contexts. Our code and data are available at https://github.com/typybench/typybench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 119 citations
- Typilus: neural type hintsMiltiadis Allamanis, Earl T. Barr, Soline Ducousso, Zheng GaoPLDI 2020 · 92 citations
- Generative Type Inference for PythonYun Peng, Chaozheng Wang, Wenxuan Wang, Cuiyun Gao et al.ASE 2023 · 29 citations
Related papers
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen et al.EMNLP 2025
- CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level TasksLingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan et al.ACL 2026
- CodeSync: Synchronizing Large Language Models with Dynamic Code Evolution at ScaleChenlong Wang, Zhaoyang Chu, Zhengxiang Cheng, Xuyi Yang et al.ICML 2025
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun et al.AAAI 2026 · 10 citations
