arxiv cs.CL

Why they matter

Feed with public evidence across Compute & infrastructure.

AI signals
730
past 30d
Sources
2
distinct domains
Discussões
0
past 30d
Latest signal
3d ago
View every signal from arxiv cs.CL →
Computer Science -- Computation and Language source: export.arxiv.org/rss/cs.CL maintainer: @tmaehara.bsky.social

Articles & links

DiffusionGemma Team, Adrien Ali Ta\"iga, James Assiene, Daniele Calandriello, Rahma Chaabouni, Jo\~ao Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, ... DiffusionGemma Technical Report https://arxiv.org/abs/2608.00146

DiffusionGemma Technical Report arxiv.org
AI Weekly's analysis
  • DiffusionGemma reportedly generates around 1,500 output tokens per second on a single NVIDIA H100 by refining blocks of 256 tokens in parallel.
  • The model is a fine-tune of the mixture-of-experts Gemma 4 base, which has 3.8B activated and 25.2B total parameters.
  • Its two-stage training pipeline of supervised denoising plus RL with sampler distillation used less than 10% of the original AR model's token budget.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 1 ↩ 0 · 4 from the directory shared this · 13d ago

Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak Characterizing Narrative Content in Web-scale LLM Pretraining Data https://arxiv.org/abs/2606.19468

Characterizing Narrative Content in Web-scale LLM Pretraining Data arxiv.org
AI Weekly's analysis
  • A new arXiv preprint introduces NarraBERT, a RoBERTa-based classifier, and applies it to 3 million passages from the 3-trillion-token Dolma corpus.
  • The framework operationalizes three narrative elements, agency, setting, and events, across 11 interpretable dimensions, trained on 400 annotated passages.
  • The authors report narrative qualities are unequally distributed across pretraining sources and topics in ways current curation practices do not measure.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 0 · 3 from the directory shared this · 59d ago

Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure https://arxiv.org/abs/2608.13545

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure arxiv.org
AI Weekly's analysis
  • LITTLECURRICULUM is an 88-billion-token pretraining corpus curated to U.S. elementary school material through Grade 5.
  • A 5-billion-parameter model trained from scratch on that corpus, LITTLELEARNER, has knowledge boundaries mapped to interpretable curriculum guidelines.
  • Post-training and in-context learning helped the model use existing knowledge but did not raise out-of-scope capabilities.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 3d ago

Zeyu Zhang, Ziliang Guo, Yihang Sun, Xichong Zhang, Xixuan Hao, Zehao Lin, Yang Zhang, Xiaoyan Zhao, Tong Shen, Bo Tang, Zhi-Qin John Xu, Junchi Yan, Haofen Wang, Xu Chen, Feiyu Xiong, Zhiyu Li, Tat-Seng Chua Metis: Memory Foundation Model https://arxiv.org/abs/2607.26760

Metis: Memory Foundation Model arxiv.org
AI Weekly's analysis
  • Metis introduces a persistent, dynamically evolving memory state built into the model backbone rather than as an external retrieval module.
  • The authors train Metis at 4B, 9B, and 27B parameters on a Qwen3.5 backbone using 8 H100 GPUs and around 406M synthesized tokens.
  • Metis-27B reports 73.77% average on the paper's constructed test set versus 1.69% for the baseline Qwen3.5-27B without context.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 19d ago

Maria Thomas, Kristina Gligoric, Nihar B. Shah Mitigating LLM-based p-Hacking by Preregistering for the Next LLM https://arxiv.org/abs/2606.27687

Mitigating LLM-based p-Hacking by Preregistering for the Next LLM arxiv.org
AI Weekly's analysis
  • A new arXiv paper proposes preregistering LLM experiments and running the confirmatory analysis on the first eligible model released after registration.
  • Across 20 models from four providers and 11 configurations, the protocol blocked p-hack transfer in 73.9% and 72.7% of cases across two tasks.
  • The authors preregistered their own experiment; of 7 configurations that hacked the prior model, 6 failed to carry over to the next.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 49d ago

Nina Begus World Wide Models: Literary Tools for Cultural AI https://arxiv.org/abs/2607.02369

[2607.02369] World Wide Models: Literary Tools for Cultural AI arxiv.org
AI Weekly's analysis
  • Nina Begus argues LLMs stage a cultural encounter that is 'massive, automated, and monolingual,' and proposes literary scholarship as the corrective toolkit.
  • The essay proposes applying world literature methods — macrostructure, circulation, and untranslatability — to build culturally literate AI.
  • The 15-page essay is forthcoming in MFS Modern Fiction Studies in 2027 and connects critical theory to structural monolingualism in AI.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 45d ago

Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context https://arxiv.org/abs/2606.26493

Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context arxiv.org
AI Weekly's analysis
  • NVIDIA's Nemotron-TwoTower splits an LM into a frozen autoregressive context tower and a trainable diffusion denoiser with bidirectional block attention.
  • The system is built on Nemotron-3-Nano-30B-A3B, a 30B hybrid Mamba-Transformer MoE backbone, and trained on roughly 2.1 trillion tokens.
  • The authors report retaining 98.7% of the autoregressive baseline's quality while delivering 2.42x higher wall-clock generation throughput.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 52d ago

Liancheng Gong, Zhiyang Wang, Yiwei Xu, Julia Mendelsohn Persuasion Index: A Theory-Guided Framework for Persuasion Analysis https://arxiv.org/abs/2606.14580

Persuasion Index: A Theory-Guided Framework for Persuasion Analysis arxiv.org
AI Weekly's analysis
  • Persuasion Index is a taxonomy of 15 dimensions grounded in psychology and communication, implemented with 55 sub-features from lexicons and rule-based detectors.
  • The authors evaluate PI on four public datasets varying in domain, style, and outcome measures, and report linear models carry meaningful predictive signal while staying lightweight.
  • PI is released as an open-source package and web interface for principled and auditable analysis of human and AI-mediated communication.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 63d ago

In arxiv cs.CL's orbit

Center = arxiv cs.CL. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.