Agentic coding is genuinely useful now, and there are some impressive reports of AI agents doing science. But how well and how reliably can they handle tasks scientists actually want to hand off, ones that bottleneck progress? How do we even measure that?? New paper🧵 arxiv.org…
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline arxiv.org
AI Weekly's analysis
→
- Coding agents solved individual stages of a fly optogenetics pipeline but could not complete the full end-to-end discovery run.
- Agents struggle most when there is no predefined criterion to iterate on and must use their own scientific judgment, a key open challenge.
- The study flags challenges absent from standard benchmarks: computational resource management and generalization to large held-out datasets.
Read full analysis →
View on Bluesky ·
♥ 16
↻ 7
↩ 1
·
2 from the directory shared this ·
48d ago
Huge credit to Kai Horstmann, who did the majority of this work across Kristin Branson and Jennifer Sun's lab, with help from Ethan Lin and Alice Robie. Paper: arxiv.org/abs/2606.07718 Task environments: github.com/kaihorstmann/neuro-d2d-eval Data: huggingface.co/datasets/kaih…
kaihorstmann/flyopto-d2d · Datasets at Hugging Face huggingface.co
This is pretty crazy, Fable will silently harm (am I understanding that right?) ML research "building pretraining pipelines, distributed training infrastructure, or ML accelerator design" jonready.com/blog/posts/c...
If Claude Fable stops helping you, you'll never know — Jonathon Ready jonready.com
Good news! Our paper evaluating agentic AI for computational neuroscience tasks was accepted to COLM (Conference on Language Modeling)! colmweb.org/index.html This work was led by Kai Horstmann, a talented PhD student, and done in collaboration with Jennifer Sun's group at Cor…
COLM 2026 colmweb.org