Anthropic red team tests frontier AI on targeting and drones
TL;DR
- Anthropic's Frontier Red Team says its top models beat the strongest human baseline on photo geolocation, scoring 37.0 km and 47.2 km median error across 6,000 photos.
- In a simulated drone environment, Opus 5 hit stationary high-visibility targets on 80% of launches but only 20% across all nine difficulty settings.
- Anthropic's Safeguards team added new classifiers to block weapons-development requests after identifying real misuse of Claude in that domain.
Anthropic's Frontier Red Team published a report measuring how frontier models perform on tasks its authors say "historically, only a set of scarce, highly-trained human experts could do": geolocating people from photos and social media posts, and writing drone control software for simulated strikes.
The specifics are the point. On outdoor photo geolocation, "Mythos Preview and Mythos 5 beat even the strongest human baseline on median distance error, scoring 37.0 km and 47.2 km across 6,000 photos." On a simulated quadcopter, "Opus 5 strikes on 80% of its launches" against stationary high-visibility targets, and "Across all nine settings, Opus 5 hits the target on 20% of 540 launches." On the harder payload delivery task in wind, "Opus 5 is the only model that succeeds with any regularity. It hits 28% of its sorties, with a median miss distance of 3.9 meters on the payloads it releases."
The report puts these in operational language. "'Kill chains,' such as 'find, fix, track, target, engage, assess,' are end-to-end conceptual models of these engagements," it says, and its evaluations aim to "better illustrate how AI progress is changing the risk landscape across different parts of the kill chain."
Open-weights systems were not far behind on some tasks. "Kimi K3 performs about as well as the frontier on easy and medium samples," and "GLM 5.2 was nearly identical to Sonnet 5 at 31.0 km." The paper says "these evaluations also underscore the urgency of research into more robust approaches to open-weights model safety."
On the closed-weight side, Anthropic writes that its "Safeguards team implemented new classifiers to detect and block requests related to weapons development after identifying misuse of Claude in this domain." A companion note from its Threat Intelligence Team, the paper says, "includes instances of AI misuse in surveillance and conventional weapons development which show threat actors already perceiving benefit from the use of AI models." Two of the researchers we follow in our Who's Who tracker shared the report as it went up.
Shared on Bluesky by 2 AI experts
-
www.anthropic.com/research/int... hey uh,
View on Bluesky →
Originally reported by anthropic.com
Read the original article →Original headline: Measuring AI capabilities in intelligence targeting and conventional weapons