β»
Robert Hawkins reposted
β»
Robert Hawkins reposted
Eugene Vinitsky
@eugenevinitsky.bsky.social
"LLMs are just n-dimensional classifiers" has a very similar flavor to "everything is just quantum mechanical interactions." www.science.org/doi/10.1126/...
science.org
View on Bluesky β
β»
Robert Hawkins reposted
Russ Poldrack
@russpoldrack.org
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis arxiv.org/abs/2608.19902 - our group's latest work, led by Zijiao Chen, on an agentic harness for neuroimaging analysis that aims to promote analytic rigor.
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis arxiv.org
AI Weekly's analysis
β
Read full analysis β
View on Bluesky β
β»
Robert Hawkins reposted
Brenden Lake
@brendenlake.bsky.social
Excited about new work from my CCN keynote (38m): youtu.be/v3J-vJfxhOE?... We often choose between Bayesian and neural net models of behavior. But what if there's a spectrum, and combining both makes better predictions? Introducing BBT w @akjagadish.bsky.social & Guangyuan arxβ¦
More accurate behavioral predictions with hybrid Bayesian-connectionist models arxiv.org
AI Weekly's analysis
β
Read full analysis β
View on Bluesky β
β»
Robert Hawkins reposted
arxiv cs.CL
@arxiv-cs-cl.bsky.social
Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives https://arxiv.org/abs/2608.08160
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives arxiv.org
AI Weekly's analysis
β
- NCP-Bench spans 100 narrative environments derived from movie synopses, with an automated checker that scores player agent vs. narrator agent consistency turn by turn.
- The best model tested, GPT-5.2, maintains only a 42% survival rate after 20 turns of interaction under unconstrained user interventions.
- Across six state-of-the-art LLMs, fact conflicts dominate failure modes with rates running from 40% to 68%.
Read full analysis β
View on Bluesky β
β»
Robert Hawkins reposted
β»
Robert Hawkins reposted
@alessandrogifford.bsky.social
(1/2) The EEG Moments Dataset (EMD) is now public! EMD is a large-scale dataset of simultaneous EEG and eye-tracking recordings of 6 participants for 1,102 naturalistic 3-second videos (while maintaining central fixation), together with rich video metadata. arxiv.org/abs/2608.β¦
A large dataset of human EEG responses to short naturalistic videos for studying dynamic visual event processing arxiv.org
View on Bluesky β
β»
Robert Hawkins reposted
@martinhebart.bsky.social
What is the nature of universal representations in AI models, and what determines whether they emerge? Our paper accepted at #neurips2026 led by @florianmahner.bsky.social & @rothj.bsky.social addressed these questions comparing 162 vision models, with intriguing results. arxiβ¦
Characterizing Universal Object Representations Across Vision Models arxiv.org
View on Bluesky β
β»
Robert Hawkins reposted
@masonyoungblood.bsky.social
Our new preprint explores how increasing interaction between humans and artificial agents and is reshaping creativity, from the perspective of cultural evolution! @katieeeeeeeeeee.bsky.social @manuelangladatort.bsky.social @camrobjones.bsky.social @gemschedel.bsky.social arxivβ¦
Collective creativity in hybrid societies arxiv.org
View on Bluesky β
β»
Robert Hawkins reposted
@iyadrahwan.bsky.social
Introducing: Time Machine Experiments? π β³ in which participants 'travel' to the past, and interact with a mind from the year 1930 (simulated by an AI with knowledge cut-off). Preprint: arxiv.org/pdf/2609.15468
arxiv.org
View on Bluesky β