DiffusionGemma Team, Adrien Ali Ta\"iga, James Assiene, Daniele Calandriello, Rahma Chaabouni, Jo\~ao Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, ... DiffusionGemma Technical Report https://arxiv.org/abs/2608.00146
DiffusionGemma Technical Report arxiv.org
AI Weekly's analysis
→
- DiffusionGemma reportedly generates around 1,500 output tokens per second on a single NVIDIA H100 by refining blocks of 256 tokens in parallel.
- The model is a fine-tune of the mixture-of-experts Gemma 4 base, which has 3.8B activated and 25.2B total parameters.
- Its two-stage training pipeline of supervised denoising plus RL with sampler distillation used less than 10% of the original AR model's token budget.
Read full analysis →
Zeyu Zhang, Ziliang Guo, Yihang Sun, Xichong Zhang, Xixuan Hao, Zehao Lin, Yang Zhang, Xiaoyan Zhao, Tong Shen, Bo Tang, Zhi-Qin John Xu, Junchi Yan, Haofen Wang, Xu Chen, Feiyu Xiong, Zhiyu Li, Tat-Seng Chua Metis: Memory Foundation Model https://arxiv.org/abs/2607.26760
Metis: Memory Foundation Model arxiv.org
AI Weekly's analysis
→
- Metis introduces a persistent, dynamically evolving memory state built into the model backbone rather than as an external retrieval module.
- The authors train Metis at 4B, 9B, and 27B parameters on a Qwen3.5 backbone using 8 H100 GPUs and around 406M synthesized tokens.
- Metis-27B reports 73.77% average on the paper's constructed test set versus 1.69% for the baseline Qwen3.5-27B without context.
Read full analysis →
Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin, Jocelyn Shen, Blake A. Richards, Alison Gopnik, Doina Precup Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? https://arxiv.org/abs/2606.06464
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? arxiv.org
AI Weekly's analysis
→
Read full analysis →
Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context https://arxiv.org/abs/2606.26493
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context arxiv.org
AI Weekly's analysis
→
- NVIDIA's Nemotron-TwoTower splits an LM into a frozen autoregressive context tower and a trainable diffusion denoiser with bidirectional block attention.
- The system is built on Nemotron-3-Nano-30B-A3B, a 30B hybrid Mamba-Transformer MoE backbone, and trained on roughly 2.1 trillion tokens.
- The authors report retaining 98.7% of the autoregressive baseline's quality while delivering 2.42x higher wall-clock generation throughput.
Read full analysis →
Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti Language Models Need Sleep https://arxiv.org/abs/2605.26099
[2605.26099] Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference arxiv.org
Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure https://arxiv.org/abs/2608.13545
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure arxiv.org
AI Weekly's analysis
→
- LITTLECURRICULUM is an 88-billion-token pretraining corpus curated to U.S. elementary school material through Grade 5.
- A 5-billion-parameter model trained from scratch on that corpus, LITTLELEARNER, has knowledge boundaries mapped to interpretable curriculum guidelines.
- Post-training and in-context learning helped the model use existing knowledge but did not raise out-of-scope capabilities.
Read full analysis →
Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots https://arxiv.org/abs/2608.05004
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots arxiv.org
AI Weekly's analysis
→
Read full analysis →
Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen Simulating Human Memory with Language Models https://arxiv.org/abs/2605.25680
[2605.25680] Simulating Human Memory with Language Models arxiv.org