Tech & AI News
Hacker News

DiffusionGemma Technical Report

DiffusionGemma is an open-weight language model that refines 256-token blocks instead of decoding token by token, reaching about 1,500 output tokens per second on a single NVIDIA H100 GPU. It fine-tunes the 4-billion-parameter Gemma mixture-of-experts model using supervised bidirectional denoising followed by reinforcement-learning sampler distillation, using under 10% of the original training token budget. The model preserves Gemma’s thinking mode, multimodal input handling, long-context capability, and can still generate autoregressively with only slight performance degradation.