NobleBlocks
Public

Decoding Speculative Decoding

Published in Underline Science Inc. • Apr 14, 2025
NobleIDNI6P28W50R81S50
Authors:
Association for Computational Linguistics 2025
,
Saurabh Agarwal
,
Shivaram Venkataraman

Abstract

Speculative Decoding is a widely used technique to speed up inference for Large Language Models (LLMs) without sacrificing quality. When performing inference, speculative decoding uses a smaller draft model to generate speculative tokens and then uses the target LLM to verify those draft tokens. The...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!