Salvaggio: AI's 'black box' is a policy choice, not mystery
TL;DR
- Essay argues the 'black box' framing is marketing cover for undisclosed training data and system prompts, not an inherent technical mystery.
- Points to Anthropic's own Claude Opus 4 safety card, where 84% of completions in a prompted roleplay described blackmail.
- Separates interpretability, the open neural-network research problem, from transparency, meaning disclosure of training data and prompts.
The industry's 'we don't understand how it works' line is a chosen stance, not a technical fact, argues Eryk Salvaggio, a fellow at Tech Policy Press, in a new essay. He splits two ideas the industry blurs: interpretability, meaning the open research problem of what happens inside a neural network, and transparency, meaning what a company chooses to publish about its training data, system prompt and evaluation criteria.
Salvaggio builds the case from Anthropic's own safety card for Claude Opus 4, where the model was 'asked... to act as an assistant at a fictional company' and, prompted toward a corner, produced blackmail text in most runs. '84% of the resulting text completions described blackmail,' the card reports. Salvaggio reads that not as a glimpse of emergent ethics but as a model doing what it was told, generating text appropriate to the scenario. A second example has Claude processing documents for an imagined drug called Zenavex and flagging fraud, again pattern-matching to a prompt rather than reasoning.
The target of the critique is a rhetorical move the essay traces to the top of the industry. Dario Amodei is quoted saying 'people outside the field are often surprised and alarmed to learn that we do not understand how our own AI creations work.' Sam Altman is quoted with a brain analogy: 'We don't understand what's happening in your brain at a neuron-by-neuron level, and yet we know you can follow some rules.' Salvaggio argues the conflation of interpretability with disclosure serves a business purpose. 'Myth-making is a crucial aspect of the AI industry, and black boxes are woven into the stories they tell.'
His central distinction is one line: 'Interpretability is not transparency. Transparency means sharing the system prompt and the data relied upon for training.' Why that stays undone, in his reading, is incentive. 'AI as an industry doesn't grow if it cannot control the models, and it cannot control the models without controlling the data.'
Five analysts in our Who's Who tracker shared the piece. Salvaggio ends on the reframe the essay is built around: 'That's not a black box. It's just a policy decision.'
Shared on Bluesky by 5 AI experts
-
Anthropic’s response and influence on the encyclical is most evident in the claim that “even its creators don’t understand how it works.” In June, I wrote about what that actually means & why Anthropic says it. (Written …
View on Bluesky → -
“Those of us outside of the AI industry don’t know what rules they are following. That’s not a black box. It’s just a policy decision.” techpolicy.press/the-black-bo... Or: it’s a remediable black box.
View on Bluesky →
Originally reported by techpolicy.press
Read the original article →Original headline: The Black Box Myth: What the Industry Pretends Not to Know About AI | TechPolicy.Press