Lune

ICML2025Top-tier venue

Prompt-based Depth Pruning of Large Language Models

Juyun Wee, Minjae Park, Jaeho Lee

2025Year
2Top-tier citations

Abstract

Depth pruning aims to reduce the inference cost of a large language model without any hardwarespecific complications, by simply removing several less important transformer blocks. However, our empirical findings suggest that the importance of a transformer block may be highly taskdependent-a block that is crucial for a task can be removed without degrading the accuracy on another task. Based on this observation, we develop a dynamic depth pruning algorithm, coined PuD-Ding (Prompt-routed Dynamic Depth Pruning), which determines which blocks to omit from the model based on the input prompt. PuDDing operates by training a lightweight router to predict the best omission set among a set of options, where this option set has also been constructed in a data-driven manner. Empirical results on commonsense reasoning benchmarks demonstrate that PuDDing effectively accelerates the inference language models, and achieves better on-task performance than static depth pruning baselines.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers2

Ask how each one uses it

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines