Decoding Speculative Decoding
Published in Underline Science Inc. • Apr 14, 2025
NobleIDNI6P28W50R81S50
Authors:,,
Association for Computational Linguistics 2025
Saurabh Agarwal
Shivaram Venkataraman
Abstract
Speculative Decoding is a widely used technique to speed up inference for Large Language Models (LLMs) without sacrificing quality. When performing inference, speculative decoding uses a smaller draft model to generate speculative tokens and then uses the target LLM to verify those draft tokens. The...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!