Artificial Intelligence Papers

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
468
past 30d
Sources
4
distinct domains
Discussions
0
past 30d
Latest signal
5h ago
View every signal from Artificial Intelligence Papers →

Articles & links

A game theory for foundation models shows new paths to rational cooperation through similarity inference Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, … https://t.co/zQvZFwMM1v [𝚌𝚜.𝙰𝙸] ht…

A game theory for foundation models shows new paths to rational cooperation through similarity inference arxiv.org
AI Weekly's analysis
  • The paper reports that foundation model agents in stylized social dilemmas consistently converge to stable cooperation, contradicting classical predictions of mutual defection.
  • The authors introduce the 'embedded Bayesian agent,' which models an agent as part of the universe it inhabits rather than an independent decision-maker.
  • They propose 'embedded equilibrium' as a new solution concept replacing the Nash equilibrium for reasoning about modern AI agents.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 9d ago

SPIRAL: Learning to Search and Aggregate Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman https://t.co/CRBpj1Mjhk [𝚌𝚜.𝙰𝙸] https://t.co/kVEHyMHKpK

SPIRAL: Learning to Search and Aggregate arxiv.org
AI Weekly's analysis
  • SPIRAL co-trains three reasoning primitives in one RL framework: sequential chain-of-thought, parallel sampling of traces, and learned aggregation of those traces.
  • The paper reports outperforming GRPO by up to 11× scaling efficiency and 15% higher performance when all three compute primitives are scaled.
  • Training uses set reinforcement learning to make parallel traces collectively useful, plus standard RL to train the aggregation step itself.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 54d ago

Beyond expert users: agents should help users construct preferences, not just elicit them Irena Saracay, Ludwig Schmidt, Carlos Guestrin https://t.co/C6dYa9Tgwh [𝚌𝚜.𝙰𝙸] https://t.co/LTMprro9iE

Beyond expert users: agents should help users construct preferences, not just elicit them arxiv.org
AI Weekly's analysis
  • New arxiv paper argues AI agents should help non-expert users construct preferences, not assume users already know what they want.
  • The authors introduce CoShop, an interactive benchmark where no tested agent exceeded 56% accuracy after five turns of dialogue.
  • Failures came from agents' limited knowledge expansion, not from difficulty finding items once preferences were specified.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 47d ago

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models Kevin Murphy https://t.co/80KY0Z6z5i [𝚌𝚜.𝙰𝙸] https://t.co/fccLmcAbkd

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models arxiv.org
AI Weekly's analysis
  • Kevin Murphy's Model Discovery Agent uses an LLM to propose candidate causal structures, then Bayesian machinery to test them from few interventions.
  • The pipeline combines sequential Monte Carlo, simulation-based inference, and value-of-information experiment design in an M-open hypothesis setting.
  • MDA is evaluated on physics, chemistry, and a new single-neuron electrophysiology benchmark introduced alongside the paper.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 4d ago

Full-bandwidth transformer Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford https://t.co/FUDCgqgkqW [𝚌𝚜.𝙰𝙸] https://t.co/eiA0jqzYAQ

Full-bandwidth transformer arxiv.org
AI Weekly's analysis
  • A 1B-parameter 'full-bandwidth' transformer adds latent feedback and reportedly matches standard models trained on roughly 1.5x more tokens.
  • Latent feedback fuses the previous top-layer hidden state with the sampled token embedding through a gated linear unit before re-entering the stack.
  • Reported gains span validation loss, 5-shot evaluation, and math and coding generation, with shorter reasoning traces at equal or better accuracy.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 5d ago

Autodata: An agentic data scientist to create high quality synthetic data Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, … https://t.co/iSchw5CkfT [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻 𝚌𝚜.𝙻𝙶] https://t.co/fC2KJEmmyE

Autodata: An agentic data scientist to create high quality synthetic data arxiv.org
AI Weekly's analysis
  • Meta researchers introduce Autodata, a method that casts an AI agent as a data scientist iteratively generating and refining synthetic training data.
  • The practical implementation is called Agentic Self-Instruct, and meta-optimizing the data scientist agent itself produced a larger uplift than static methods.
  • On legal reasoning tasks, a 4B parameter model trained on agent-made data reportedly beat a 397B parameter baseline.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 52d ago

Bayesian control for coding agents Theodore Papamarkou, Vladislav Smirnov, Viktor Mazanov, Artem Vazhentsev, Preslav Nakov, Timothy Baldwin, Artem Shelmanov https://t.co/1EUIZ7fmTy [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻] https://t.co/5sFFzguwnn

Bayesian control for coding agents arxiv.org
AI Weekly's analysis
  • A new arxiv paper recasts coding-agent orchestration as cost-sensitive sequential hypothesis testing managed by a Bayesian controller.
  • The controller decides dynamically whether to gather more evidence, refine the solution, run a verifier, or stop the run.
  • Authors report the approach is most valuable when verification is costly and critics are informative but imperfect, across six generators and nine benchmarks.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 53d ago

auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation Ben Prystawski, Kushin Mukherjee, Daniel Wurgaft, Linas Nasvytis, Michael Y. Li, Noah D. Goodman, Michael C. Frank https://t.co/U0Bv3bc9yi [𝚌𝚜.𝙰𝙸] https://t.co/S9vC7tD3F3

auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation arxiv.org
AI Weekly's analysis
  • Auto-psych uses nested loops: an inner loop generates probabilistic cognitive models, an outer loop designs and runs online human experiments.
  • In three independent human experiments, the system's discovered theories fit the data better than theories drawn from the scientific literature.
  • The benchmark task was a classic cognitive psychology problem about how people perceive randomness in coin flip sequences.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 51d ago

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo https://t.co/UdRrMz9Tfg [𝚌𝚜.𝙰𝙸] https://t.co/I87HYLYXKD

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems arxiv.org
AI Weekly's analysis
  • TsuGO scores how LLMs organize a search on Go tsumego problems, not just whether they land on the correct final move.
  • The paper reports most models behave 'much closer to unguided search algorithms than to neural-guided KataGo' on these puzzles.
  • Longer chain-of-thought and higher token efficiency do not necessarily produce better search; stronger models commit to the right candidate earlier.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 1d ago

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li https://t.co/djg9LPOhqg [𝚌𝚜.𝙰𝙸] https://t.co/bOf1jPYR0k

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure arxiv.org
AI Weekly's analysis
  • SkillZip posts a 31.2% compression rate versus 9.2% for SkillReducer, with average task performance of 0.577 against 0.544.
  • The method avoids task rollouts and evaluation sets entirely, relying on structural analysis of the skill artifact itself.
  • Cross-model transfer on LiveMath retains 0.97 of performance versus 0.91 for the baseline, suggesting better portability across agent backbones.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 2d ago

Matryoshka Language Model Suites Nathan Godey, Yoav Artzi https://t.co/a8Er3e7gMX [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻] https://t.co/6lTyfpvIFy

Matryoshka Language Model Suites arxiv.org
AI Weekly's analysis
  • Nathan Godey and Yoav Artzi train sub-models of 500M, 1.5B, and 3B parameters nested inside a single architecture instead of separately.
  • The paper claims 36% less training compute versus independently trained models while matching baseline benchmarks, validation metrics, and out-of-domain perplexities.
  • Speculative decoding throughput reportedly rises 14 to 26% because the smaller draft model is already embedded inside the larger verifier.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 4d ago

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo https://t.co/6Hf1t8qPOE [𝚌𝚜.𝙰𝙸] https://t.co/BzvnO…

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents arxiv.org
AI Weekly's analysis
  • BCSD, from 11 authors led by Tianjun Pan, targets the gap between finishing an agent task and truly using provided natural-language skills.
  • It scores each trajectory twice, once with added Meta-Skill guidance and once with general guidance pruned, then rescales RL advantage at the token level.
  • The authors report strongest overall performance across model scales on ALFWorld and WebShop, though the abstract does not publish specific baseline gains.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 2 from the directory shared this · 4d ago