shipfeedAI news, curated daily

01:36:42 CET
24 AUG01:36:42shipfeed
pull to refreshlast sync
Just in — 30 new
§ models · storyline

Google's DiffusionGemma proves you don't need from-scratch training

Google DeepMind converts Gemma 4 into a diffusion model using less than 10 percent training budget with parallel generation at 1,500 tokens per second.

Aug 9 · · primary fetch1 sourceupdated Aug 9 ·

Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second.

Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks. The article Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model appeared first on The Decoder.

read full article on the-decoder.com
§ sources1 publication · timeline below
  1. the-decoder.comGoogle's DiffusionGemma proves you don't need to train from scratch to build a text diffusion modelprimary