Hacker News
vLLM v0.28.0
vLLM 0.28.0 adds Decode Context Parallel, fused FlashKDA kernels, GEMM-RS sequence parallelism, and shared-expert sharding that reduces GPU memory use by ~17 GiB, plus ROCm support. It introduces sparse MLA for DeepSeek V4, DFlash2 and DSpark speculative decoding, weight offloading, multi-layer MTP KV caches, attention-free models, tiered KV-cache disk offloading, a Rust gRPC frontend with multimodal image inference, higher default token limits, and new models such as Muse Glimmer, Qwen 3.8 on ROCm, and vision-tower LoRA for Gemma 4.