Tech & AI News
Hacker News

Speculative Decoding in vLLM on AMD GPUs

Speculative decoding in vLLM introduces a draft-and-verify step where a lightweight draft model proposes multiple tokens and the target model validates them in one pass, allowing batch commitment of output tokens. Benchmarks on AMD Instinct MI300X and MI355X GPUs with ROCm show throughput varies by drafting method, proposal length, model family, draft checkpoint, workload, and acceptance rate. Five methods—native MTP, Gemma-4 MTP, EAGLE-3, DFlash, and DSpark—are compared, with implementation details, tuning advice, and observability guidance included.