science.org web signal

Astrophysicists Openly Debate Whether AI Is Doing Their Science

TL;DR

  • In September 2025, a Flatiron Institute system called Denario ran live during an NYU physics talk, generating research ideas, analyses and near-complete papers.
  • About one in ten Denario outputs raises an intriguing question, and the system has fabricated data in some cases, keeping human review essential.
  • NYU's David Hogg has posted an arXiv paper, 'Why do we do astrophysics?', arguing against both unrestricted AI adoption and outright bans.

Something interesting is happening in a field you would not expect to be first to the microphone on this. Astrophysics has started openly asking whether AI systems are doing its work, and whether that is still the same work. Science reported on the debate playing out at Harvard, the Flatiron Institute and NYU, and the specific incident that set it off is worth sitting with.

In September 2025, a guest speaker at the NYU physics department gave a talk while a system called Denario ran live in the background, generating entire scientific projects as he spoke. Denario, built at the Flatiron Institute, scours the literature, produces research ideas, executes analyses and drafts near-complete papers. According to the reporting, a seventh-year grad student, Matthew Daunt, walked out and told the NYU and Flatiron astrophysicist David Hogg that he was 'not cattle', a response to the guest speaker's line that 'you don't need grad students anymore'.

That reaction has hardened into a real discussion about what astrophysics is for. Hogg has since posted an arXiv paper, 'Why do we do astrophysics?', which argues against both the unrestricted-adoption and outright-ban ends of the policy spectrum. Meanwhile the tools keep landing wins the field cannot ignore. Alyssa Goodman's Harvard group is credited with producing 'the single best map of spiral arm kinematics ever, by a factor of 100' using AI-assisted data fitting, and Cecilia Garraffo's AstroAI group at the Center for Astrophysics is collaborating with Anthropic and Google DeepMind on discipline-specific systems.

The honest caveat is that Denario itself is not the end-state anyone is arguing about. Only about one in ten of its outputs yields an interesting question, and it has fabricated data in some cases, so human review remains essential. Ethan Vishniac of the American Astronomical Society is quoted worrying that the sheer quantity of low-quality output could 'strangle the system' of peer review. What the reporting does not fully answer is how authorship and credit get assigned when an agent writes the paper, or how fast the same argument will surface in biology, chemistry and materials science, which are already on Denario's test list.

The thing worth watching is not the tools; it is that a small, prestigious field is having the conversation out loud rather than quietly adopting. That is the useful move for any leader whose team is about to be handed a tool that does the middle chunk of the job.

Shared on Bluesky by 3 AI experts