AI Firehose

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
466
past 30d
Sources
1
distinct domains
Discusiones
1
past 30d
Latest signal
1d ago
View every signal from AI Firehose →
Daily-updated stream of AI research from ArXiv

Articles & links

A study shows the "Matthew Effect" in reinforcement learning for large language models, where RL boosts easy tasks while neglecting harder ones. The "Never Give Up" method reallocates compute resources to enhance success on hard tasks, promising progress in AI. https://arxiv.o…

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up arxiv.org
AI Weekly's analysis →
  • The paper argues RL post-training gives large gains on easy problems but small gains on hard ones, calling the pattern the Matthew Effect.
  • Never Give Up (NGU) keeps sampling a given problem until a correct answer appears, reallocating compute toward harder items via asynchronous RL.
  • On the Deepscaler math benchmark and the Manufactoria coding task, the authors say NGU improves performance per compute, especially on harder problems.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 0 · 3 from the directory shared this · 10d ago

Brain Researcher enhances neuroimaging analysis by embedding methodological judgment, achieving a 70.2% tool selection increase. This innovation reshapes scientific claims into auditable processes, improving the credibility of neuroimaging research. https://arxiv.org/abs/2608.…

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis arxiv.org
AI Weekly's analysis →
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 4 from the directory shared this · 36d ago

Superintelligent AI, designed through a solipsistic lens, risks failing at cooperation due to undermining behaviors from interactions among adaptive agents. This challenges paradigms and calls for cooperative systems emphasizing human agency and institutional design. https://a…

Solipsistic Superintelligence is Unlikely to be Cooperative arxiv.org
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 116d ago

Cognitive science is set for a breakthrough with AI integration, allowing generalizable models of cognition via naturalistic tasks. This method reshapes intelligence understanding, yielding insights and hypotheses about human cognition with complex data. https://arxiv.org/abs/…

[2502.20349] Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior arxiv.org
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 125d ago

Stanford's Spiral framework redefines language model training by merging sequential, parallel, and aggregative inference, boosting reasoning efficiency up to 15% over previous methods. https://arxiv.org/abs/2606.23595

SPIRAL: Learning to Search and Aggregate arxiv.org
AI Weekly's analysis →
  • SPIRAL co-trains three reasoning primitives in one RL framework: sequential chain-of-thought, parallel sampling of traces, and learned aggregation of those traces.
  • The paper reports outperforming GRPO by up to 11× scaling efficiency and 15% higher performance when all three compute primitives are scaled.
  • Training uses set reinforcement learning to make parallel traces collectively useful, plus standard RL to train the aggregation step itself.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 95d ago

FreeToken transforms personal machines into edge-native inference platforms, enabling efficient serving of MoE models up to 753B parameters. This innovation narrows the accessibility gap for frontier AI, making cutting-edge capabilities practical for individuals. https://arxiv…

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution arxiv.org
AI Weekly's analysis →
  • FreeToken's abstract claims the system serves a 753B GLM-5.2 mixture-of-experts model on a single workstation GPU.
  • The same system is claimed to run a 284B model on a gaming desktop and a 35B model on an 8GB laptop GPU.
  • Author list includes Shuo Yang, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu and Ion Stoica.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 0 · 3 from the directory shared this · 38d ago

Researchers and the New Jersey Public Defender's office teamed up to create an AI retrieval tool that boosts legal research, using realistic data and innovative query techniques. This enhances advocacy efficiency and sets a precedent for AI in public interest law. https://arxi…

Legal Retrieval for Public Defenders arxiv.org
AI Weekly's analysis →
  • A team partnered with the New Jersey Office of the Public Defender to build NJ BriefBank, a tool that surfaces relevant appellate briefs.
  • The paper reports that existing retrieval benchmarks fail to transfer to real public defense research.
  • Adding domain knowledge such as query expansion with legal reasoning, domain-specific data, and curated synthetic examples improved retrieval quality.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 46d ago

D-OPSD transforms training for step-distilled diffusion models, enabling on-policy self-distillation to learn new concepts without sacrificing efficient few-step inference. This enhances image quality and response speed for AI-generated content. https://arxiv.org/abs/2605.05204

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models arxiv.org
AI Weekly's analysis →
  • The paper argues ordinary supervised fine-tuning of step-distilled diffusion models compromises their inherent few-step inference capability.
  • D-OPSD treats the model as both teacher, seeing text plus target-image information, and student, seeing only text features.
  • The authors claim their approach lets models learn new concepts and styles without sacrificing original few-step capacity.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 121d ago

Researchers devised a statistical model to estimate uncertainty dynamics in text generation by smoothing noisy data from large language models. This advancement reduces resampling costs, enhancing insights into LLM reasoning and decisions in complex text generation. https://ar…

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation arxiv.org
AI Weekly's analysis →
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 35d ago

Prime Agent enhances long-horizon agency in AI using a self-improving harness that boosts language model capabilities through persistent execution and real-time collaboration with recursive subagents, significantly improving performance on coding and reasoning tasks. https://a…

Prime Agent: A Self-Improving RLM Harness arxiv.org
AI Weekly's analysis →
  • Prime Agent, an open-source harness from Prime Intellect, raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5%, per the arXiv abstract posted 24 August 2026.
  • The design pairs a persistent IPython REPL following the Recursive Language Model abstraction with a Continual Harness that preserves histories, memories, skills, prompts and subagent specifications across trajectories.
  • The paper reports parity or better than native and popular harnesses on long-context coding, GPU-kernel generation, emulator construction and autonomous nanoGPT speedruns.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 32d ago

Terminal-Universe transforms agent trajectories into reusable environments, generating 37.3k task-rich settings for enhanced coding and agent training. This approach outperforms traditional methods, providing scalable solutions in coding. https://arxiv.org/abs/2609.04148

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments arxiv.org
AI Weekly's analysis →
  • Terminal-Universe reconstructs code-agent training workspaces by replaying file operations logged in prior trajectories, producing 37.3k task-sufficient environments from public traces.
  • SFT on Qwen3.5-27B using the generated corpus lifted Terminal-Bench 2.1 single-round scores by 11.9 points and EvoCode-Bench v2 MT@4 multi-round by 13.8.
  • A completion agent fills in missing files and dependencies after the replay, and the framework then synthesizes cross-codebase queries and multi-round user sessions.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 22d ago

In AI Firehose's orbit

Center = AI Firehose. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you AI Firehose? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/ai-firehose-column-social)