Link to tech report: https://t.co/g710KPO6UR
- DiffusionGemma reportedly generates around 1,500 output tokens per second on a single NVIDIA H100 by refining blocks of 256 tokens in parallel.
- The model is a fine-tune of the mixture-of-experts Gemma 4 base, which has 3.8B activated and 25.2B total parameters.
- Its two-stage training pipeline of supervised denoising plus RL with sampler distillation used less than 10% of the original AR model's token budget.