DiffusionGemma Team, Adrien Ali Ta\"iga, James Assiene, Daniele Calandriello, Rahma Chaabouni, Jo\~ao Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, ... DiffusionGemma Technical Report https://arxiv.org/abs/2608.00146
DiffusionGemma Technical Report arxiv.org
AI Weekly's analysis
→
- DiffusionGemma reportedly generates around 1,500 output tokens per second on a single NVIDIA H100 by refining blocks of 256 tokens in parallel.
- The model is a fine-tune of the mixture-of-experts Gemma 4 base, which has 3.8B activated and 25.2B total parameters.
- Its two-stage training pipeline of supervised denoising plus RL with sampler distillation used less than 10% of the original AR model's token budget.
Read full analysis →
Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure https://arxiv.org/abs/2608.13545
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure arxiv.org
AI Weekly's analysis
→
- LITTLECURRICULUM is an 88-billion-token pretraining corpus curated to U.S. elementary school material through Grade 5.
- A 5-billion-parameter model trained from scratch on that corpus, LITTLELEARNER, has knowledge boundaries mapped to interpretable curriculum guidelines.
- Post-training and in-context learning helped the model use existing knowledge but did not raise out-of-scope capabilities.
Read full analysis →
Zeyu Zhang, Ziliang Guo, Yihang Sun, Xichong Zhang, Xixuan Hao, Zehao Lin, Yang Zhang, Xiaoyan Zhao, Tong Shen, Bo Tang, Zhi-Qin John Xu, Junchi Yan, Haofen Wang, Xu Chen, Feiyu Xiong, Zhiyu Li, Tat-Seng Chua Metis: Memory Foundation Model https://arxiv.org/abs/2607.26760
Metis: Memory Foundation Model arxiv.org
AI Weekly's analysis
→
- Metis introduces a persistent, dynamically evolving memory state built into the model backbone rather than as an external retrieval module.
- The authors train Metis at 4B, 9B, and 27B parameters on a Qwen3.5 backbone using 8 H100 GPUs and around 406M synthesized tokens.
- Metis-27B reports 73.77% average on the paper's constructed test set versus 1.69% for the baseline Qwen3.5-27B without context.
Read full analysis →
Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin, Jocelyn Shen, Blake A. Richards, Alison Gopnik, Doina Precup Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? https://arxiv.org/abs/2606.06464
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? arxiv.org
AI Weekly's analysis
→
Read full analysis →
Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context https://arxiv.org/abs/2606.26493
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context arxiv.org
AI Weekly's analysis
→
- NVIDIA's Nemotron-TwoTower splits an LM into a frozen autoregressive context tower and a trainable diffusion denoiser with bidirectional block attention.
- The system is built on Nemotron-3-Nano-30B-A3B, a 30B hybrid Mamba-Transformer MoE backbone, and trained on roughly 2.1 trillion tokens.
- The authors report retaining 98.7% of the autoregressive baseline's quality while delivering 2.42x higher wall-clock generation throughput.
Read full analysis →
Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti Language Models Need Sleep https://arxiv.org/abs/2605.26099
[2605.26099] Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference arxiv.org
Jared Moore, Andrea Mock, Yifan Mai, Jacy Reese Anthis, Ryan Louie, William Agnew, Ashish Mehta, Kevin Klyman, Percy Liang, Nick Haber, Eric Lin, Desmond C. Ong DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots https://arxiv.org/abs/2608.05004
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots arxiv.org
AI Weekly's analysis
→
Read full analysis →
Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen Simulating Human Memory with Language Models https://arxiv.org/abs/2605.25680
[2605.25680] Simulating Human Memory with Language Models arxiv.org