Artificial Intelligence Papers

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
834
past 30d
Sources
5
distinct domains
Discusiones
0
past 30d
Latest signal
57m ago
View every signal from Artificial Intelligence Papers →

Articles & links

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis Zijiao Chen, Nicholas Lu, Xinhui Li, Jocelyn A. Ricard, Ce Ju, Huan H. Wang, Christian Kindermann, … https://t.co/7cdp7kbZWF [𝚌𝚜.𝙰𝙸 𝚚-𝚋𝚒𝚘.𝙽𝙲] https://t.co/lGnlyqztds

Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis arxiv.org
AI Weekly's analysis
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 4 from the directory shared this · 3d ago

A game theory for foundation models shows new paths to rational cooperation through similarity inference Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, … https://t.co/zQvZFwMM1v [𝚌𝚜.𝙰𝙸] ht…

A game theory for foundation models shows new paths to rational cooperation through similarity inference arxiv.org
AI Weekly's analysis
  • The paper reports that foundation model agents in stylized social dilemmas consistently converge to stable cooperation, contradicting classical predictions of mutual defection.
  • The authors introduce the 'embedded Bayesian agent,' which models an agent as part of the universe it inhabits rather than an independent decision-maker.
  • They propose 'embedded equilibrium' as a new solution concept replacing the Nash equilibrium for reasoning about modern AI agents.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 17d ago

Full-bandwidth transformer Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford https://t.co/FUDCgqgkqW [𝚌𝚜.𝙰𝙸] https://t.co/eiA0jqzYAQ

Full-bandwidth transformer arxiv.org
AI Weekly's analysis
  • A 1B-parameter 'full-bandwidth' transformer adds latent feedback and reportedly matches standard models trained on roughly 1.5x more tokens.
  • Latent feedback fuses the previous top-layer hidden state with the sampled token embedding through a gated linear unit before re-entering the stack.
  • Reported gains span validation loss, 5-shot evaluation, and math and coding generation, with shorter reasoning traces at equal or better accuracy.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 13d ago

SPIRAL: Learning to Search and Aggregate Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman https://t.co/CRBpj1Mjhk [𝚌𝚜.𝙰𝙸] https://t.co/kVEHyMHKpK

SPIRAL: Learning to Search and Aggregate arxiv.org
AI Weekly's analysis
  • SPIRAL co-trains three reasoning primitives in one RL framework: sequential chain-of-thought, parallel sampling of traces, and learned aggregation of those traces.
  • The paper reports outperforming GRPO by up to 11× scaling efficiency and 15% higher performance when all three compute primitives are scaled.
  • Training uses set reinforcement learning to make parallel traces collectively useful, plus standard RL to train the aggregation step itself.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 62d ago

Beyond expert users: agents should help users construct preferences, not just elicit them Irena Saracay, Ludwig Schmidt, Carlos Guestrin https://t.co/C6dYa9Tgwh [𝚌𝚜.𝙰𝙸] https://t.co/LTMprro9iE

Beyond expert users: agents should help users construct preferences, not just elicit them arxiv.org
AI Weekly's analysis
  • New arxiv paper argues AI agents should help non-expert users construct preferences, not assume users already know what they want.
  • The authors introduce CoShop, an interactive benchmark where no tested agent exceeded 56% accuracy after five turns of dialogue.
  • Failures came from agents' limited knowledge expansion, not from difficulty finding items once preferences were specified.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 55d ago

Terminal Agents: A Survey of AI Agents in Command-Line Environments Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen https://t.co/wDmCLGosQt [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝚂𝙴] https://t.co/4Xli3X2hdW

Terminal Agents: A Survey of AI Agents in Command-Line Environments arxiv.org
AI Weekly's analysis
  • A new arXiv survey groups AI "terminal agents" around a common lens: systems whose main loop is mediated by command execution and textual feedback.
  • The authors introduce a seven-dimensional terminal competence profile linking system architecture, competence acquisition, and evaluation.
  • They argue prevailing benchmarks emphasize final outcomes and expose process quality, recovery, and governance unevenly.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 1d ago

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding Giuseppe Destefanis, Tomaso Aste https://t.co/JeBsDN6L0r [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝚂𝙴] https://t.co/acBWUJM25J

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding arxiv.org
AI Weekly's analysis
  • Across 1,902 runs, a new instrument represents each multi-agent coding run as a temporal network of agents, files, messages, reads and writes.
  • Shared files can replace direct messaging, cutting output tokens by about 42% at eight agents on message-heavy work.
  • In a sealed replication with marked placeholder files, agents still tried to reach hidden grading material in four fifths of 244 runs.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 5d ago

ASI-Bench: At the Dawn of Artificial Superintelligence Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, … https://t.co/FKR1HT5XiH [𝚌𝚜.𝙰𝙸] https://t.co…

ASI-Bench: At the Dawn of Artificial Superintelligence arxiv.org
AI Weekly's analysis
  • On ASI-Bench, the average score across 18 frontier agent-model configurations falls from 50.91 with full guidance to 26.62 when the agent picks the method.
  • The benchmark covers 60 project-level research tasks across 11 scientific domains, built by more than 40 experts and over 31,000 human hours.
  • Every task goes through expert review, AI-assisted auditing, sandbox execution, and scorer validation before it is admitted.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 5d ago

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, … https://t.co/EkebOwiVk8 [𝚌𝚜.𝙰𝙸] https://t.co/…

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence arxiv.org
AI Weekly's analysis
  • Apodex Discovery selected 20 problems from 423 candidates surveyed across 561 industries and 16 sectors for its initial benchmark release.
  • On AAV capsid design the system beat the published state of the art by 7% across viability, tropism, structure prediction, and generative design.
  • The HDS6 rubric grades Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of whether the final task succeeded.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 10d ago

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models Kevin Murphy https://t.co/80KY0Z6z5i [𝚌𝚜.𝙰𝙸] https://t.co/fccLmcAbkd

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models arxiv.org
AI Weekly's analysis
  • Kevin Murphy's Model Discovery Agent uses an LLM to propose candidate causal structures, then Bayesian machinery to test them from few interventions.
  • The pipeline combines sequential Monte Carlo, simulation-based inference, and value-of-information experiment design in an M-open hypothesis setting.
  • MDA is evaluated on physics, chemistry, and a new single-neuron electrophysiology benchmark introduced alongside the paper.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 12d ago

Autodata: An agentic data scientist to create high quality synthetic data Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, … https://t.co/iSchw5CkfT [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻 𝚌𝚜.𝙻𝙶] https://t.co/fC2KJEmmyE

Autodata: An agentic data scientist to create high quality synthetic data arxiv.org
AI Weekly's analysis
  • Meta researchers introduce Autodata, a method that casts an AI agent as a data scientist iteratively generating and refining synthetic training data.
  • The practical implementation is called Agentic Self-Instruct, and meta-optimizing the data scientist agent itself produced a larger uplift than static methods.
  • On legal reasoning tasks, a 4B parameter model trained on agent-made data reportedly beat a 397B parameter baseline.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 60d ago

Bayesian control for coding agents Theodore Papamarkou, Vladislav Smirnov, Viktor Mazanov, Artem Vazhentsev, Preslav Nakov, Timothy Baldwin, Artem Shelmanov https://t.co/1EUIZ7fmTy [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻] https://t.co/5sFFzguwnn

Bayesian control for coding agents arxiv.org
AI Weekly's analysis
  • A new arxiv paper recasts coding-agent orchestration as cost-sensitive sequential hypothesis testing managed by a Bayesian controller.
  • The controller decides dynamically whether to gather more evidence, refine the solution, run a verifier, or stop the run.
  • Authors report the approach is most valuable when verification is costly and critics are informative but imperfect, across six generators and nine benchmarks.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 61d ago