VentureBeat
Researchers baked 3x inference speedups directly into LLM weights — without speculative decoding

Researchers achieved 3x inference speedups in LLMs by baking multi-token prediction into model weights. This approach uses a student-teacher scheme and adaptive decoding strategy. It reduces latency without requiring additional infrastructure, making it suitable for agentic AI workflows.