used Claude Opus 5 to chain two OpenAI vulnerabilities—a libheif memory bug reachable via HEIF image uploads on the Discourse-based OpenAI community forum—to take over multiple OpenAI employee accounts and access OpenAI's internal monorepo in authorized bug bounty research
Reported: successfully opened a pull request in OpenAI's internal monorepo; OpenAI paid a $6,500 bug bounty and fixed the flaws
Ran Gemini in a capture-the-flag security evaluation conducted by Israeli firm Irregular; a misconfiguration allowed the model to reach the real internet
Reported: Gemini autonomously hacked three companies: brute-forced passwords at one, scraped credentials from public repos at two others; stopped hacking on its own in all three cases
Strix used an AI agent to discover an unauthenticated Harbor registry and a live GITHUB_TOKEN with admin access to Baseten's repositories baked into a 2023 Docker image
Reported: AI agent found admin and push access to Baseten's main product repo, its GitOps deployment repo, and its Homebrew distribution channel
Launched SafeMind dual-model agentic system where Red Tempest probes for attack paths and Blue Solano patches them in a closed loop on a digital twin of the customer environment, shipping inside Falcon with standalone access via Project QuiltWorks
Implemented chain-of-thought monitoring on all RL training and evaluations involving tools for GPT-5.6 Sol-class and higher models and all inference on the Astra model
Reported: monitoring overhead at roughly 20% of the inference compute being monitored
Used Wiz Red Agent AI to audit Snowflake's snowflake-connector-net repository and discover a shell-injection vulnerability introduced by a GitHub Copilot Autofix patch
Reported: Red Agent identified an exploitable shell-injection flaw; unauthenticated attacker exfiltrated the Jira token for [email protected] within five days, granting read access to engineering, security compliance and bug bounty projects
deployed GPT-5.6-Cyber through the Daybreak Red program for offensive security vulnerability research, discovering previously unknown Chrome V8 vulnerabilities
Reported: model discovered two previously unknown V8 vulnerabilities in Chrome that can be chained to corrupt memory and bypass the V8 heap sandbox, patched by Google under CVE-2026-15903; model responds to 95% of sensitive security queries versus 57.3% for predecessor GPT-5.5-Cyber
deployed autonomous AI research system (HTTP Terminator) to test 30,000 candidate desync vectors against thousands of authorized websites to identify novel HTTP desync vulnerabilities
Reported: identified roughly 700 vulnerable targets including banks, government infrastructure, security products, and an airport; generated new attack classes including dual-matching Content-Length pattern, dangling-byte technique, and shared-parser confusion; exposed an Apache Traffic Server zero-day
Google uses AI tools to discover, validate, triage and fix Chrome vulnerabilities and is piloting twice-weekly security releases to absorb the resulting volume.
Reported: Chrome versions 149 and 150 fixed 1,072 bugs between them.
Searchlight Cyber used GPT-5.6 Sol Ultra for autonomous multi-agent analysis of the default WordPress codebase.
Reported: About 10 hours and $25 of model usage found a pre-authenticated SQL-injection-to-remote-code-execution chain affecting more than 500 million WordPress instances.
OpenAI uses the internal GPT-Red model to automate red teaming and adversarially train frontier models against prompt injection.
Reported: Prompt-injection attacks fell from more than 90% success on GPT-5 to under 23% on GPT-5.6; GPT-Red reached 84% attack success versus 13% for human red-teamers on a benchmark.
Physical-security company Verkada is adopting Nvidia's Cosmos world foundation models and Physical AI Data Factory toolkit to scale AI across its 2.4 million connected devices.
Reported: The integration has already delivered a 68% accuracy gain in spatial-temporal video search.
Under Project Glasswing, Cloudflare deployed Anthropic's Mythos model against real-world cyber frontier threat scenarios across Cloudflare infrastructure, publishing model behavior, capability boundaries, and detection findings.
creating a standalone AI mission organization as part of the agency's largest restructuring in at least a decade, with units expected to begin operations in mid-October
Installed more than 3,000 AI-powered license plate reader cameras via grants to cities and counties since 2023 for vehicle surveillance
Reported: Texas Gov. Abbott ordered state agencies to pause funding amid privacy concerns and misuse reports; cities and counties canceling contracts; Lufkin officer facing 100 counts for surveilling 11 people
Agents wore Meta's AI smart glasses during immigration enforcement operations across six states; internal memo now bars employees from wearing the devices due to risk of capturing sensitive data
The Pentagon awarded Perennial Autonomy a $500M-ceiling JIATF-401 contract for AI drone-on-drone counter-drone interceptors, making counter-drone autonomy a discrete DoD program of record.
SoftBank launched an OpenAI-powered 'Patching as a Service' providing vulnerability assessment, remediation planning and implementation advisory to the top 3,000 companies behind Japan's critical infrastructure, after internally validating the approach by running OpenAI's models across its own systems.
Every entry names the organisation and links its source. Outcome figures are quoted as reported, never estimated. Vendor announcements without a named customer are excluded. Halted and reversed deployments are kept on purpose.
We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site.
Privacy policy