Lune

NeurIPS2025Top-tier venue

AutoJudge: Judge Decoding Without Manual Annotation

Roman Garipov, Fedor Velikonivtsev, Ivan Ermakov, Ruslan Svirschevski, Vage Egiazarian, Max Ryabinin

2025Year
14Citations
3Top-tier citations

Abstract

We introduce AutoJudge 1 , a method that accelerates large language model (LLM) inference with task-specific lossy speculative decoding. Instead of matching the original model output distribution token-by-token, we identify the generated tokens that affect the downstream quality of the response, relaxing the distribution match guarantee so that the "unimportant" tokens can be generated faster. Our approach relies on a semi-greedy search algorithm to test which of the mismatches between target and draft models should be corrected to preserve quality and which ones may be skipped. We then train a lightweight classifier based on existing LLM embeddings to predict, at inference time, which mismatching tokens can be safely accepted without compromising the final answer quality. We evaluate AutoJudge with multiple draft/target model pairs on mathematical reasoning and programming benchmarks, achieving significant speedups at the cost of a minor accuracy reduction. Notably, on GSM8K with the Llama 3.1 70B target model, our approach achieves up to ≈2× speedup over speculative decoding at the cost of a ≤1% drop in accuracy. When applied to the LiveCodeBench benchmark, AutoJudge automatically detects programming-specific important tokens, accepting ≥25 tokens per speculation cycle at a 2% drop in Pass@1. Our approach requires no human annotation and is easy to integrate with modern LLM inference frameworks.

  • For greedy decoding, it checks that the drafted tokens are the same as the target model's own next token predictions. For sampling, it uses a procedure that matches the sampling probabilities [Leviathan et al., 2023].

  • For notation simplicity, we assume that the GENERATE(•, •) function can be called with a prefix of a response. In that case, we assume that the total response length (and not just newly generated tokens) does not exceed Tmax, so that the response cannot grow indefinitely with each subsequent replacement.

  • More precisely, find the fastest-to-generate sequence, accounting for the differences in response length.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9bddc110-a3d8-4627-9cd4-e664acd0d46f

Cited by top-tier papers3

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines