Tech & AI News
Hacker News

DFlash 2: Keep Drafting Parallel

DFlash 2 extends parallel drafting by generating over 20% more tokens per verification pass with only ~1% extra latency, yielding 16–25% throughput gains across benchmarks; the Qwen3.8-27B drafter now runs 2.7–3.4× faster than autoregressive decoding at batch size 1. The system keeps the top 16 candidates per position, scores adjacent pairs using logit scores and 256-dimensional embeddings, and selects coherent token paths without sequential correction.