AccueilAI Use-Case Library › AI in cybersecurity: 15 real deployments

AI in cybersecurity: 15 real deployments

Vulnerability discovery, red teaming, incident response and the attackers using the same tools.

15deployments
13in production or with results
8with a reported outcome
0halted or reversed
Aug 30, 2026last updated

Security is the function where frontier models changed the work fastest, on both sides. The entries cover autonomous vulnerability discovery, patching services, red-team agents and documented offensive use by state-linked groups.

Software & Tech 12 deployments

Anthropic

using Alice to red-team and monitor production AI systems for jailbreaks, prompt injection, and agentic misuse

In production Aug 25, 2026 Source: techstartups.com
Wiz

Used Wiz Red Agent AI to audit Snowflake's snowflake-connector-net repository and discover a shell-injection vulnerability introduced by a GitHub Copilot Autofix patch

Reported: Red Agent identified an exploitable shell-injection flaw; unauthenticated attacker exfiltrated the Jira token for [email protected] within five days, granting read access to engineering, security compliance and bug bounty projects

Results reported Aug 17, 2026 Wiz Red Agent Source: wiz.io
OpenAI

Using GPT-5.6 Sol Ultrafast mode internally for incident-response log analysis

In production Aug 14, 2026 GPT-5.6 Sol Source: helpnetsecurity.com
OpenAI

deployed GPT-5.6-Cyber through the Daybreak Red program for offensive security vulnerability research, discovering previously unknown Chrome V8 vulnerabilities

Reported: model discovered two previously unknown V8 vulnerabilities in Chrome that can be chained to corrupt memory and bypass the V8 heap sandbox, patched by Google under CVE-2026-15903; model responds to 95% of sensitive security queries versus 57.3% for predecessor GPT-5.5-Cyber

Results reported Aug 10, 2026 GPT-5.6-Cyber Source: the-decoder.com
PortSwigger

deployed autonomous AI research system (HTTP Terminator) to test 30,000 candidate desync vectors against thousands of authorized websites to identify novel HTTP desync vulnerabilities

Reported: identified roughly 700 vulnerable targets including banks, government infrastructure, security products, and an airport; generated new attack classes including dual-matching Content-Length pattern, dangling-byte technique, and shared-parser confusion; exposed an Apache Traffic Server zero-day

Results reported Aug 7, 2026 Source: portswigger.net
Google

Google uses AI tools to discover, validate, triage and fix Chrome vulnerabilities and is piloting twice-weekly security releases to absorb the resulting volume.

Reported: Chrome versions 149 and 150 fixed 1,072 bugs between them.

Results reported Jul 30, 2026 Source: Wired
XBOW

XBOW used its autonomous offensive-security agent to test Microsoft Bing Images and disclose two unauthenticated command-injection flaws.

Reported: The agent found CVE-2026-32194 and CVE-2026-32191, both rated CVSS 9.8; Microsoft fixed both before disclosure.

Results reported Jul 24, 2026 XBOW autonomous security agent Source: The Hacker News
Searchlight Cyber

Searchlight Cyber used GPT-5.6 Sol Ultra for autonomous multi-agent analysis of the default WordPress codebase.

Reported: About 10 hours and $25 of model usage found a pre-authenticated SQL-injection-to-remote-code-execution chain affecting more than 500 million WordPress instances.

Results reported Jul 20, 2026 GPT-5.6 Sol Ultra Source: Searchlight Cyber
OpenAI

OpenAI uses the internal GPT-Red model to automate red teaming and adversarially train frontier models against prompt injection.

Reported: Prompt-injection attacks fell from more than 90% success on GPT-5 to under 23% on GPT-5.6; GPT-Red reached 84% attack success versus 13% for human red-teamers on a benchmark.

Results reported Jul 15, 2026 GPT-Red Source: MIT Technology Review
Microsoft

Microsoft is reorganizing its cybersecurity business around AI-assisted vulnerability discovery and automated threat response.

Reported: The restructuring replaced senior executives and cut hundreds of roles.

Results reported Jul 14, 2026 Security Copilot Source: The Information
Microsoft

Microsoft uses the multi-model MDASH agentic scanning harness to discover Windows vulnerabilities faster.

In production Jul 10, 2026 MDASH Source: The Register
Cloudflare

Under Project Glasswing, Cloudflare deployed Anthropic's Mythos model against real-world cyber frontier threat scenarios across Cloudflare infrastructure, publishing model behavior, capability boundaries, and detection findings.

Pilot May 18, 2026 Claude Mythos Source: Cloudflare

Military 1 deployment

Grimfengxi

Used DeepSeek to generate exploit code for cyberattacks

In production Aug 24, 2026 DeepSeek Source: bloomberg.com

Government & Public Sector 1 deployment

US Treasury, CISA, DHS and Department of Defense

The Gold Eagle federal clearinghouse uses frontier AI to scan government and private-sector systems for vulnerabilities and prioritize patches.

In production Jul 14, 2026 Gold Eagle Source: CyberScoop

Telecom 1 deployment

SoftBank

SoftBank launched an OpenAI-powered 'Patching as a Service' providing vulnerability assessment, remediation planning and implementation advisory to the top 3,000 companies behind Japan's critical infrastructure, after internally validating the approach by running OpenAI's models across its own systems.

Announced Jun 16, 2026 Source: SoftBank

Every entry names the organisation and links its source. Outcome figures are quoted as reported, never estimated. Vendor announcements without a named customer are excluded. Halted and reversed deployments are kept on purpose.