uses Raindrop Simulations to replay production AI agent traces against proposed code changes using synthetic copies of databases, payment APIs, and communications tools to catch hallucinations, tool misuse, and behavior drift in pull requests
Home › AI Use-Case Library › AI in software development: 28 real deployments
AI in software development: 28 real deployments
Named deployments in software development, grouped by industry.
Software & Tech 25 deployments
uses Raindrop Simulations to replay production AI agent traces against proposed code changes using synthetic copies of databases, payment APIs, and communications tools to catch hallucinations, tool misuse, and behavior drift in pull requests
uses Raindrop Simulations to replay production AI agent traces against proposed code changes using synthetic copies of databases, payment APIs, and communications tools to catch hallucinations, tool misuse, and behavior drift in pull requests
Deployed a GLM-5.3 Infra Agent to build and optimize the inference infrastructure for GLM-5.3-Flash on a cluster of more than 100,000 Chinese-made AI accelerators
Reported: End-to-end throughput tripled from baseline in under two weeks; per-token cost and hardware efficiency described as comparable to mainstream Nvidia GPUs
Google opened Claude Opus 5 access to all engineers through its internal Antigravity IDE for software development, replacing a prior policy requiring Gemini use
using Temporal Durable Execution platform for AI agent workloads
using Temporal Durable Execution platform to keep long-running AI agents fault-tolerant
Reported: OpenAI's Temporal usage grew 60-fold in under a year
Launched Project HydraFusion multi-model orchestration system routing coding tasks across multiple LLMs using direct-solve, cascade, and critique-and-revise patterns, available as research preview in Copilot CLI
Reported: 4.9 pt gains on TerminalBench 2.1 at 67% lower cost vs Claude Opus 5; 36% cost reduction on DeepSWE; 65% lower cost on CheckpointBench
Graduated Muse Code terminal AI coding agent from beta to production, adding inter-session messaging over Unix sockets, multi-subagent workflow orchestration, and a TypeScript SDK
Used AI agents to author approximately 2,000 pull requests to build DoltLite Beta, a SQLite fork with Git-style versioning
Reported: DoltLite Beta shipped, built via roughly 2,000 AI-agent-authored pull requests
AI writes approximately 70% of Grindr's code across engineering operations
Reported: Engineering output has 2.5x'd since July 2025
AI coding assistant using Anthropic, Google, and SpaceXAI models in production; OpenAI model supply being wound down effective November 12 following SpaceX acquisition
Reported: OpenAI represents only ~5% of Cursor's traffic
Consolidating Trae coding platform and Coze agent-building tool into Doubao super-app, and planning to launch Doubao Work productivity agent
Routing production inference traffic for Hermes Agent through Ox Alpha model
Uses Factories, version-controlled pipelines that move tickets through spec, implementation, review and verification with coding agents, to handle internal tasks
Reported: factories already handle 30-35% of Warp's own internal tasks
Built and uses Berd internally, a desktop app for managing AI agents, files, skills and sessions across the Goose framework
Used GitHub Copilot Autofix to generate a security patch for snowflake-connector-net, which replaced a safe input pattern with raw string interpolation of a GitHub issue title
Reported: Autofix-generated patch introduced an exploitable shell-injection vulnerability; unauthenticated attacker exfiltrated Jira token for [email protected] within five days of the patch
Integrated xAI's Grok 4.6 reasoning model into GitHub Copilot for agentic coding and multi-step workflows
Integrated Nvidia NeMo Switchyard router into Devin Desktop to reduce AI inference cost
Reported: cut mean cost 28%
deployed AI Spend Console to map per-employee and per-team AI token usage against productivity signals and route spend across multiple AI providers to reduce costs
Reported: token spend dropped from 40% to 15% of R&D headcount budget; July costs were 37% of April despite similar 600B token volumes
Databricks routes AI coding tasks by complexity, defaults away from frontier models when cheaper models clear the bar, and trims coding-harness prompt overhead to control enterprise coding-agent costs.
Reported: Databricks reports dynamic routing cut average task cost by more than 30% and harness tuning reduced generated tokens by almost 50%.
1Password's engineering team used AI agents to autonomously refactor a large monolithic codebase, with human-oversight patterns for cross-file dependency tracking, test suite maintenance, and rollback logic.
Reported: Team reports meaningful velocity gains while flagging specific failure modes
EPAM is building a practice of 10,000 Claude-certified architects, including 250 forward-deployed engineer 'Black Belts', to deliver enterprise AI for Global 2000 clients using Claude models, Claude Code, and the Claude Agent SDK.
Reported: 1,300 architects already certified; 5,000 targeted by end of Q3 2026
LlamaIndex generates the overwhelming majority of its codebase with AI, per CEO Jerry Liu.
Reported: Roughly 95% of the company's codebase is now AI-generated
Google generates most of its new code with AI, with AI-written code human-reviewed, per Pichai's disclosure at Cloud Next.
Reported: AI-written, human-reviewed code went from ~30% (April 2025) to 50% (last fall) to 75% today, with one complex migration completed 6x faster year-over-year
Aerospace & Defense 1 deployment
Acquired Cursor to build AI coding capabilities inside SpaceXAI
Government & Public Sector 1 deployment
Alberta deployed AI agents to scan provincial software and rebuild government services across 27 ministries.
Reported: The government says about 50 agents scanned 466 million lines of code in roughly 20 hours and reduced a five-month subsidy-portal rebuild to four or five days at about 95% lower cost.
Transportation 1 deployment
Uber rolled out Claude Code across its 5,000-engineer organization, with adoption accelerated by an internal leaderboard ranking teams by AI usage volume.
Reported: Claude Code adoption jumped from 32% to 84%; monthly API costs run $500-$2,000 per engineer; the entire 2026 AI tools budget was exhausted by April
Every entry names the organisation and links its source. Outcome figures are quoted as reported, never estimated. Vendor announcements without a named customer are excluded. Halted and reversed deployments are kept on purpose.