How LLMs Get Faster Without Changing Their Outputs | Speculative Decoding
1 directory member surfaced this signal.
“Speculative Decoding is the coolest trick for speeding up LLM inference! Check out the video and learn why rejection sampling preserves quality, and how methods such as draft trees, Medusa, MTP, EAGLE, and DFlash further accelerate LLM inference. youtu.be/l…”