Yes, I am quoted in this. But more generally this is the kind of detailed indepth reporting on AI we need so much more of, reporting that doesn't stop at "I saw this error in a response" but contextualizes when and where such errors occur against a baseline. www.npr.org/2026/0…
Mike Caulfield
Researcher with public evidence across AI research.
- AI signals
- 4 past 30d
- Sources
- 4 distinct domains
- Discussions
- 15 past 30d
- Latest signal
- 1d ago
Articles & links
My bizarre prediction is since no one can tell smart glasses from real glasses in plastic frames that people are going to migrate to frameless and thin metal glasses just to avoid looking like they are perving on people. New style trend. www.theguardian.com/commentisfre...
I've wondered a long time why LLMs are so bad at writing prompts for users to use when they seem to be so good at prompting on the fly. It turns out the two things are related.
I think this post is one of my more important ones if you're allowed to say things like that. I hope it's clear? mikecaulfield.substack.com/p/ai-as-fuzz...
Claude created a mess with a dumb way it tends to write prompts. I wrote a blog post about this failure mode, but that still left me with a massive mess to tidy up. Then I came up with an idea: have Fable read my post & fix the issue. It fucking worked. mikecaulfield.substack.…
These are short teachable habits and if we spent a fraction of the time we spend yelling at one another about the latest Futurism post actually teaching students how to use AI like this the world would be a better place. More on four habits of prompting here open.substack.com/…
www.theatlantic.com/technology/a...
I did a bit of Fable (and Opus) assisted research on a difference I noticed in how Fable makes decisions on tagging data. The upshot is that Fable seems to treat "certainty" thresholds differently but predictably. This is an AI report here, but could be useful to some claude.a…
Recent commentary
The key skills people are going to need to use LLMs effectively are around evaluation and verification. I think many people see that as "Oh, so I have to be the LLM's copy editor now" but this is because most people don't understand evaluation and verification in any systemic way.
I don't think people fully understand the level of bullshit they consume on this site about AI daily. The new story about people cancelling Anthropic bc of the watermark is based on a reporter finding a dozen people on X saying that, there is no drop (yet) at least.
The fact that AI models still can't deal with the conflation problem of film facts is a bit frustrating. No, the Agatha Christie film Endless Night, which has in it the actress who plays Miss Moneypenny in the Bond films, is not a Bond film, a fact easily verifiable against IMDB, Wikidata, etc.
In fairness the reason people think hallucinations are still the core problem of AI is hallucinations are absolutely still a problem in the *free* models. This one about Soderbergh's noir The Underneath is wild.
I know my questions aren't the norm but I swear I get hallucinations (old style ones!) a quarter of the time I ask AI overview a film question. I was looking up this joke...
these facts appeared in an AI summary and I thought, come on
In Mike Caulfield's orbit
Center = Mike Caulfield. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Mike Caulfield? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/mikecaulfield-bsky-social)