Leiden Declaration AI and mathematics
16 experts across 5 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Clément Canonne
“A list of principles put forth by mathematicians, for mathematicians and other researchers, regarding the use of AI in research. "Number #9 will surprise you!" leidendeclaration.ai”
evidence ↗
16 experts
5 communities
1 sources clustered
“A list of principles put forth by mathematicians, for mathematicians and other researchers, regarding the use of AI in research. "Number #9 will surprise you!" leidendeclaration.ai”
“Also worth re-upping: Mathematicians standing up against how their field is being used as marketing for OpenAI and the like: leidendeclaration.ai >>”
2 experts discussed this · 5 posts
Emily M. Bender: Also worth re-upping: Mathematicians standing up against how their field is being used as marketing for OpenAI and the like: leidendeclaration.ai >>
Emily M. Bender: See also: www.scientificamerican.com/article/open...
Olivia Guest: Also worth re-upping: Mathematicians standing up against how their field is being used as marketing for OpenAI and the like: leidendeclaration.ai >>
Open the full discussion →
11 experts across 5 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
4 attributable expert contributions
· Tim Kellogg, Boris Power, René Walter
“OpenAI has an automated AI “research intern” openai.com/index/resear...”
evidence ↗
11 experts
5 communities
1 sources clustered
Research & technical analysis
4 experts
Evidence, methods and technical implications.
“OpenAI has an automated AI “research intern” openai.com/index/resear...”
Building & implementation
2 experts
How teams are shipping and applying it.
“The impact of AI-native development at OpenAI • Researchers use $600+/day of AI tokens with the top 10% at $7,000+ • Humans still plan, but OpenAI says it hit “automated research intern” in 2026 and targets an automated researcher by 2028. • The need for in…”
Anthropic AI misuse September report
15 experts across 4 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Michiel Bakker
“@kotekjedi_ml Here Ant's report https://t.co/22vHruWVqH And the paper from @kotekjedi_ml @DavidSchmotz @iliaishacked https://t.co/wuCpNZaQ0d”
evidence ↗
15 experts
4 communities
1 sources clustered
Concern & critique
5 experts
Risks, limits and unintended consequences.
“will have an article about ai agents some point next week but where i'm at rn is we are seeing AI get access to accounts/systems/the internet en masse. as directed by humans, that is gonna cause big problems anthropic threat report yesterday was interesting…”
Building & implementation
1 expert
How teams are shipping and applying it.
“Thomson Reuters announced it is moving off Claude to Alibaba's Qwen to cut costs. From today's Anthropic report: Alibaba extracted 151M+ Claude exchanges to help train Qwen. TR now deploys Westlaw on a vast trove of stolen US IP. @AnthropicAI's report: http…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“@kotekjedi_ml Here Ant's report https://t.co/22vHruWVqH And the paper from @kotekjedi_ml @DavidSchmotz @iliaishacked https://t.co/wuCpNZaQ0d”
8 experts across 4 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Shannon Mattern
““Anthropic said it could not determine whether the research served a legitimate or nefarious purpose because valid biological inquiry… can also help engineer dangerous pathogens.” Reminiscing on the days when we had peer review panels with enough subject ex…”
evidence ↗
8 experts
4 communities
1 sources clustered
Building & implementation
2 experts
How teams are shipping and applying it.
“"Anthropic said it had disrupted several potential plots this year by scientists who used its leading artificial intelligence models to conduct research that could have helped develop biological weapons."”
Policy & governance
1 expert
Rules, institutions and accountability.
“Madness that we're not regulating this. Anthropic is just deciding on a case-by-case basis whether a military institute means well when it works on increasing virus capabilities. www.nytimes.com/2026/09/10/u...”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
““Anthropic said it could not determine whether the research served a legitimate or nefarious purpose because valid biological inquiry… can also help engineer dangerous pathogens.” Reminiscing on the days when we had peer review panels with enough subject ex…”
Tao AI math misalignment essay
9 experts across 5 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
2 attributable expert contributions
· Nathan Lambert, Justin Hendrix
“A great read. I have similar feelings about how AI labs approach progress directly and without nurturing of scientific communities & intuition. The math research community went through the transition the fastest, so it was felt most. Other fields next. terr…”
evidence ↗
9 experts
5 communities
1 sources clustered
Research & technical analysis
2 experts
Evidence, methods and technical implications.
“"...the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community."”
Concern & critique
1 expert
Risks, limits and unintended consequences.
“If rewarding AI mathematicians by the same metric that we have rewarded humans ones make math go badly, then we are probably rewarding the human ones badly. With good reward metrics, we shouldn't care who wins them. https://t.co/CO6Z9o3mcW”
3 experts across 3 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Berna Devezer
“ivy full prof who publishes 25 papers a year calls for academics to publish less”
evidence ↗
3 experts
3 communities
1 sources clustered
Policy & governance
1 expert
Rules, institutions and accountability.
“the idea these corporations will imppse voluntary constraints on their pursuit of wealth and power in a country too corrupt to have working regulators is a hoot”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“ivy full prof who publishes 25 papers a year calls for academics to publish less”
Opportunity & adoption
1 expert
New capabilities, benefits and practical upside.
“tech bros promise to slow down the money printing machine again www.washingtonpost.com/technology/2...”
New
AI field signal
Signal
1h ago
3 experts across 2 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
2 attributable expert contributions
· Mark Riedl, Jonathan Stray
“I've found it instructive watching is happening in the physical sciences with LLM-accelerated research. This passage from "The Last Astronomers" in Science resonated with me www.science.org/content/arti... (here is the arXiv article referenced: arxiv.org/ab…”
evidence ↗
3 experts
2 communities
1 sources clustered
Research & technical analysis
2 experts
Evidence, methods and technical implications.
“I've found it instructive watching is happening in the physical sciences with LLM-accelerated research. This passage from "The Last Astronomers" in Science resonated with me www.science.org/content/arti... (here is the arXiv article referenced: arxiv.org/ab…”
4 experts across 3 network communities independently surfaced this.
Why this matches
Research & technical analysis reaction
2 attributable expert contributions
· Alison Gopnik, Frank Pasquale
“So pleased that our paper on empowerment and causal models is out and freely available as part of this impressive special issue on world models and AI, with Melanie Mitchell, Josh Tenenbaum, Tom Griffiths and many other stars. royalsocietypublishing.org/rst…”
evidence ↗
4 experts
3 communities
1 sources clustered
Research & technical analysis
2 experts
Evidence, methods and technical implications.
““While LLMs can generate fluent linguistic output and perform well on benchmark tasks, such performance does not entail human-like linguistic competence, and many evaluation standards incorrectly equate task success with understanding.” royalsocietypublishi…”
New
AI field signal
Signal
43m ago
2 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Julian Togelius
“A few years ago, Georgios Yannakakis and I wrote an article about what to do as an AI researcher when a giant AI company steamrolls your research field. I just looked at it again and, well, it’s still relevant. ieeexplore.ieee.org/document/104...”
evidence ↗
2 experts
2 communities
1 sources clustered
“A few years ago, Georgios Yannakakis and I wrote an article about what to do as an AI researcher when a giant AI company steamrolls your research field. I just looked at it again and, well, it’s still relevant. ieeexplore.ieee.org/document/104...”
Sakana AI product page
3 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Sakana AI
“Peak performance across hard benchmarks: • Best on 5/8 benchmarks (DeepSWE, Chartography, Toolathon, GDP.pdf, SWEFish) • Chartography: 48.3 (outperforming Opus 5 & Fable 5) • DeepSWE: 74.3 Achieved without Fable 5, Fable 5.1, or GPT-6-Astra in the agent poo…”
evidence ↗
3 experts
1 community
1 sources clustered
Building & implementation
1 expert
How teams are shipping and applying it.
“By orchestrating the world's models, we are building the resilient infrastructure required for AI sovereignty. Try: sakana.ai/fugu Blog: sakana.ai/fugu-max-rel... 🐡”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“Peak performance across hard benchmarks: • Best on 5/8 benchmarks (DeepSWE, Chartography, Toolathon, GDP.pdf, SWEFish) • Chartography: 48.3 (outperforming Opus 5 & Fable 5) • DeepSWE: 74.3 Achieved without Fable 5, Fable 5.1, or GPT-6-Astra in the agent poo…”
LLM war judgment alignment paper
2 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· John Horton
“Interesting paper from @maxchupilkin illustrating a kind of 'Volkswagen emissions test" effect where LLMs respond differently about war when being told then are being tested for alignment https://t.co/rMWObcyPZI 1/ https://t.co/g9rysSZk3p”
evidence ↗
2 experts
0 communities
1 sources clustered
“Interesting paper from @maxchupilkin illustrating a kind of 'Volkswagen emissions test" effect where LLMs respond differently about war when being told then are being tested for alignment https://t.co/rMWObcyPZI 1/ https://t.co/g9rysSZk3p”
“Language models judge war differently when tested for alignment Maxim Chupilkin https://t.co/YfyoMfeA08 [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝚈] https://t.co/0voM3qtf8M”
clinical LLM deterministic math solver
2 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· AI Firehose
“MIT researchers introduced a method to enhance clinical calculators, allowing models to produce Python code for calculations. The 'Program-Solve' approach yields better accuracy in medical computations, surpassing standard arithmetic. https://arxiv.org/abs/…”
evidence ↗
2 experts
1 community
1 sources clustered
“MIT researchers introduced a method to enhance clinical calculators, allowing models to produce Python code for calculations. The 'Program-Solve' approach yields better accuracy in medical computations, surpassing standard arithmetic. https://arxiv.org/abs/…”
“Towards a Deterministic Math Solver for Clinical Language Models Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi https://t.co/aAT7Z0DMoe […”
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Christopher Manning
“Holistic Evaluation of Language Models – https://t.co/jEje8m0mDp ELEPHANT: Measuring and understanding social sycophancy in LLMs – https://t.co/sHYjxabKQx Sycophantic AI decreases prosocial intentions and promotes dependence – https://t.co/aTSrtInvKS AI gen…”
evidence ↗
1 expert
1 community
1 sources clustered
“Holistic Evaluation of Language Models – https://t.co/jEje8m0mDp ELEPHANT: Measuring and understanding social sycophancy in LLMs – https://t.co/sHYjxabKQx Sycophantic AI decreases prosocial intentions and promotes dependence – https://t.co/aTSrtInvKS AI gen…”
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Christopher Manning
“Holistic Evaluation of Language Models – https://t.co/jEje8m0mDp ELEPHANT: Measuring and understanding social sycophancy in LLMs – https://t.co/sHYjxabKQx Sycophantic AI decreases prosocial intentions and promotes dependence – https://t.co/aTSrtInvKS AI gen…”
evidence ↗
1 expert
1 community
1 sources clustered
“Holistic Evaluation of Language Models – https://t.co/jEje8m0mDp ELEPHANT: Measuring and understanding social sycophancy in LLMs – https://t.co/sHYjxabKQx Sycophantic AI decreases prosocial intentions and promotes dependence – https://t.co/aTSrtInvKS AI gen…”
Fugu Ultra v2 model release
2 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Sakana AI
“Fugu Ultra v2 is now live on OpenRouter! 🐡 openrouter.ai/sakana/fugu-... Our flagship orchestration engine built for peak performance on complex multi-step reasoning, autonomous research, and full-stack software development.”
evidence ↗
2 experts
1 community
1 sources clustered
“Fugu Ultra v2 is now live on OpenRouter! 🐡 openrouter.ai/sakana/fugu-... Our flagship orchestration engine built for peak performance on complex multi-step reasoning, autonomous research, and full-stack software development.”
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Deb Raji
“Fwiw academia has always played a critical role in third party accountability & AI auditing - challenges abound but the commitment to publication and peer review also incentivizes some degree of rigor & public engagement (we write a bit about that h…”
evidence ↗
1 expert
1 community
1 sources clustered
“Fwiw academia has always played a critical role in third party accountability & AI auditing - challenges abound but the commitment to publication and peer review also incentivizes some degree of rigor & public engagement (we write a bit about that h…”
New
AI field signal
Signal
4h ago
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Deb Raji
“I feel like we're due a new version of this paper for the "frontier AI" audit ecosystem: https://t.co/vTfyyQJk7r”
evidence ↗
1 expert
1 community
1 sources clustered
“I feel like we're due a new version of this paper for the "frontier AI" audit ecosystem: https://t.co/vTfyyQJk7r”
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Alexander Doria
“Yes with CoT. Transformer circuit also has experiments with straight (simpler) operations. transformer-circuits.pub/2025/attribu...”
evidence ↗
1 expert
1 community
1 sources clustered
“Yes with CoT. Transformer circuit also has experiments with straight (simpler) operations. transformer-circuits.pub/2025/attribu...”