HarmProfile finds harm scales with capability across 23 LLMs
TL;DR
- HarmProfile compiles over 80,000 validated harmful outputs from 23 frontier LLMs across 13 model families.
- The paper finds both harmfulness and diversity of failures grow with model capability.
- Artifacts are organized into 15 harm categories and 57 subcategories, with source code released on GitHub.
A new benchmark logs more than 80,000 validated harmful outputs from 23 frontier large language models and reports that both the harmfulness and diversity of failures grow with model capability. The corpus, called HarmProfile, spans 13 model families and sorts artifacts into 15 harm categories and 57 subcategories, according to the paper on arXiv.
The authors frame the outputs themselves as the object of study, not evidence in an attack report. 'Model risk can be characterized from the content, severity, and variation of its safety failures,' they write, drawing an analogy to how linguistic behavior can be characterized from an utterance corpus.
The empirical claim runs against a common industry line. Frontier LLMs 'reliably produce harmful content at scale, yet exhibit distinct risk profiles,' the paper says, and 'both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface.'
The public abstract does not publish per-model scores, does not name which providers rank worst, and does not describe how the 80,000 artifacts were validated. Source code is posted at fresh-ma/HarmProfile.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: HarmProfile: 80K Harmful Outputs From 23 Frontier LLMs Show Harmfulness and Diversity Grow With Capability