Melanie Mitchell

Professor at the Santa Fe Institute, AI and cognitive science

Why they matter

Professor at the Santa Fe Institute, AI and cognitive science with public evidence across AI research.

AI signals
2
past 30d
Sources
2
distinct domains
Discussions
6
past 30d
Latest signal
16d ago
View every signal from Melanie Mitchell →
Professor, Santa Fe Institute. Research on AI, cognitive science, and complex systems. Website: https://melaniemitchell.me Substack: https://aiguide.substack.com/

Articles & links

Question for Twitter hivemind: Attached is a sentence from the Hugging Face initial report on hacking incident (https://t.co/PuIHBdLCXK). Can anyone explain to me what "executing many thousands of individual actions across a swarm of short-lived sandboxes" means? https://t.co/…

Security incident disclosure — July 2026 huggingface.co
AI Weekly's analysis
  • The attacker's agent ran 17,000+ actions across short-lived sandboxes, compressing multi-stage lateral movement into a single weekend.
  • Commercial model APIs blocked forensic requests containing real exploit artifacts, forcing Hugging Face to pivot to open-weight GLM 5.2 on private infrastructure.
  • The intrusion entered via a remote dataset RCE loader and configuration template injection, not through the model-serving layer.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 8 from the directory shared this · 16d ago

Worth revisiting this prescient paper from @mmitchell.bsky.social and others from @hf.co : arxiv.org/abs/2502.02649

Fully Autonomous AI Agents Should Not be Developed arxiv.org
AI Weekly's analysis
  • Four Hugging Face researchers argue in an arXiv paper that fully autonomous AI agents should not be developed.
  • The paper proposes a five-level autonomy scale, from a Level 1 simple processor to a Level 5 agent that writes and executes its own code.
  • Their central claim is that risks to people scale with how much control the user cedes to the agent.
Read full analysis →
View on Bluesky · ♥ 74 ↻ 16 ↩ 3 · 2 from the directory shared this · 29d ago

Yet another way to jailbreak LLMs: When asking them for forbidden information, make them think the request is part of their chain-of-thought: arxiv.org/abs/2603.12277

Prompt Injection as Role Confusion arxiv.org
AI Weekly's analysis
  • A new ICML 2026 paper argues prompt injection succeeds because LLMs judge the source of text by how it sounds, not by role tags.
  • The authors' role probes show injected text lands in the same representational space as the trusted role it imitates, predicting attack success before generation.
  • Their CoT Forgery attack hits around 60% success on frontier models, while a destyling defense reportedly drops success from 61% to roughly 10%.
Read full analysis →
View on Bluesky · ♥ 67 ↻ 18 ↩ 1 · 2 from the directory shared this · 38d ago

I'm just now reading @randomwalker.bsky.social's writeup of his ICML keynote. It is very insightful. Everyone thinking about the future of work in the age of AI should read this. www.normaltech.ai/p/what-will-...

What will be left for us to work on? normaltech.ai
View on Bluesky · ♥ 88 ↻ 19 ↩ 3 · 6 from the directory shared this · 36d ago

I wrote a piece for the Yale Review on AI and "jagged intelligence". (Note: headline was not written by me.) yalereview.org/article/mela...

Melanie Mitchell: The Dangerous Unknowns at the Heart of LLMs yalereview.org
AI Weekly's analysis
  • AI systems excel on some tasks while failing surprisingly on similar ones, a pattern Mitchell calls 'jagged intelligence.'
  • Ilya Sutskever, OpenAI cofounder, is quoted saying LLMs 'generalize dramatically worse than people' in a way that is 'very fundamental.'
  • Mitchell says AI researchers, herself included, are still struggling to design effective evaluation methods for LLMs' uneven capabilities.
Read full analysis →
View on Bluesky · ♥ 126 ↻ 36 ↩ 8 · 5 from the directory shared this · 91d ago

Also, Shannon's original proposals for language modeling? And even earlier, Markov chains? In case it's helpful, I wrote a bit about the history here: oecs.mit.edu/pub/zp5n8ivs...

oecs.mit.edu
View on Bluesky · ♥ 11 ↻ 0 ↩ 0 · 54d ago

Recent commentary

The ultimate tldr on the OpenAI "rogue model" hacking incident 😅 (h/t @linege1 on X) But it would be useful to know the prompts OpenAI gave to the model....

View on Bluesky · ♥ 136 ↻ 47 ↩ 8 · 46d ago

Why is Google AI Overview still so incredibly bad??

View on Bluesky · ♥ 39 ↻ 2 ↩ 12 · 69d ago

Who wants to read yet another think piece on the OpenAI / Hugging Face incident? Choices: (1) Stop, enough already! (2) I still have questions (3) Yes, please write one! Let me know your choice, and if you still have questions, what are they?

View on Bluesky · ♥ 34 ↻ 0 ↩ 15 · 5d ago

In Melanie Mitchell's orbit

Center = Melanie Mitchell. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Melanie Mitchell? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/melanie-mitchell)