Thomas Dietterich
Safe and robust AI/ML, former AAAI president
Safe and robust AI/ML, former AAAI president with public evidence across AI research.
- AI signals
- 3 past 30d
- Sources
- 3 distinct domains
- Discussões
- 17 past 30d
- Latest signal
- 19d ago
Articles & links
That said, assessment of the significance of the work or the depth of insight are things I suspect LLMs will not do well. See also openreview.net/pdf?id=xz1gO... "Position: AI Should Verify, Not Judge, Scientific Work" by Prabhant Singh, et al. and arxiv.org/abs/2511.21843
That said, assessment of the significance of the work or the depth of insight are things I suspect LLMs will not do well. See also openreview.net/pdf?id=xz1gO... "Position: AI Should Verify, Not Judge, Scientific Work" by Prabhant Singh, et al. and arxiv.org/abs/2511.21843
- FLAWS is a new benchmark of 713 paper-error pairs built by using LLMs to insert claim-invalidating errors into peer-reviewed papers.
- GPT-5 led five frontier models with 39.1% identification accuracy at k=10, meaning it missed the planted flaw more often than not.
- Claude Sonnet 4.5, DeepSeek Reasoner v3.1, Gemini 2.5 Pro and Grok 4 were also evaluated, all below GPT-5's top score.
Great post by Daphne Koller on linkedin: www.linkedin.com/posts/daphne...
Recent commentary
At @arxiv.bsky.social, we are receiving a new type of paper that I call an "I did this experiment" paper. These papers typically report some experiment with an LLM or LLM "agentic" workflow. They are the kind of experiments an "insider" engineer would run to optimize a system. 1/
In Thomas Dietterich's orbit
Center = Thomas Dietterich. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.