Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dongyan Lin, Jocelyn Shen, Blake A. Richards, Alison Gopnik, Doina Precup Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? https://arxiv.org/abs/2606.06464
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? arxiv.org
AI Weekly's analysis
→
Read full analysis →
Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context https://arxiv.org/abs/2606.26493
Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context arxiv.org
AI Weekly's analysis
→
- NVIDIA's Nemotron-TwoTower splits an LM into a frozen autoregressive context tower and a trainable diffusion denoiser with bidirectional block attention.
- The system is built on Nemotron-3-Nano-30B-A3B, a 30B hybrid Mamba-Transformer MoE backbone, and trained on roughly 2.1 trillion tokens.
- The authors report retaining 98.7% of the autoregressive baseline's quality while delivering 2.42x higher wall-clock generation throughput.
Read full analysis →
Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti Language Models Need Sleep https://arxiv.org/abs/2605.26099
[2605.26099] Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference arxiv.org
Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen Simulating Human Memory with Language Models https://arxiv.org/abs/2605.25680
[2605.25680] Simulating Human Memory with Language Models arxiv.org
Joseph Marvin Imperial, Junhong Liang, Belal Shoer, Abdullah Barayan, Rodrigo Wilkens, Omar Mussa, Dawn Knight, Eug\'enio Ribeiro, Ekaterina Kochmar, ... ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation https://arxiv.org/abs/2606.05421
ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation arxiv.org
AI Weekly's analysis
→
- ComplexityMT benchmarks machine translation across Arabic, Dutch, English, French, Hindi and Russian using CEFR as the measure of text complexity.
- Higher CEFR levels make texts more difficult to translate, and MT systems shift the CEFR level of the target versus the source for most languages.
- The study evaluates three open-weight models, one closed model and one commercial machine translation system on two CEFR-grounded tasks.
Read full analysis →
Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao, Fan Zhang, Shijue Huang, Haitao Mi, Zhenwen Liang, Tianqing Fang, Yi R. Fung DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment https://arxiv.org/abs/2607.07820
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment arxiv.org
AI Weekly's analysis
→
- A 9B-parameter web agent reaches 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA using self-distillation.
- The DeepSearch-World environment supplies 420,000 multi-hop QA tasks built from entity-level random walks with reproducible search and page-reading.
- The training loop cycles through trajectory generation, filtering, data mixing, and fine-tuning without relying on a larger teacher model.
Read full analysis →