AI Atlas

Daily updates on real-worldAI deployments worldwide.

← All posts
Article · 1 Oct 2026

AI in Software: The Best-Measured Deployments Are on the Industry's Own Engineering Floor

Industry Deep DiveSoftwareCode ReviewAI AgentsDeveloper ProductivityCustomer SupportMulti-Agent Systems
Share

At a glance

  • AI Atlas holds 115 published use cases in Software from 88 companies in 21 countries. 111 are deployments, 2 are experiments and 2 are research. The United States accounts for 64 of them, China for 9.
  • The best-measured cases are software companies using AI on their own engineering work: code review, CI on-call and security triage. These cases report before-and-after numbers that vendor product pages almost never give.
  • Code review is the first place where AI stopped being a suggestion and became a gate. Companies split on what that gate may do on its own. One lets the agent approve a fifth of its pull requests; another blocks merges that fail a trust threshold.
  • The builders arrived at the same architecture independently: an orchestrator that routes work to narrow sub-agents, with context engineering doing more of the work than model choice.
  • Outside engineering and support the evidence thins fast. Talent and hiring platforms and regional language models make up a visible share of the set, but only a few of them report operating results.
  • At least 18 of the 115 cases describe AI supplied by a software company but used in another field, from robotaxis to classroom tutoring. This article sets those aside.

I. Code review became the first gate AI was allowed to hold

The pressure started upstream. Once code generation got faster, human review became the bottleneck. At Atlassian the median time from pull request to merge had crept past three days, and engineers waited an average of 18 hours for a first comment. Atlassian made its agent the automated first reviewer on every pull request in early 2025. Atlassian: Rovo Dev as first reviewer on every PR reports that this cut median cycle time by 45%, took the wait for a first comment to zero, and got new engineers to their first merge five days sooner. A separate year-long evaluation across more than 1,900 internal repositories is the most rigorous result in the set. It was accepted at ICSE 2026, and in it Atlassian: Rovo Dev Code Reviewer online evaluation measured a 30.8% cut in cycle time. It also found that 38.7% of the AI's comments led to a code change, against 44.5% for human comments. The agent's comments are nearly as actionable as a colleague's, not more so.

Where the companies part ways is authority.

  • The agent may approve. In April 2026, Intercom: PR review agent that auto-approves reported that 93% of its pull requests were agent-driven and more than 19% were approved with no human reviewer. Intercom says AI-authored backend changes are reverted 0.53% of the time against 5.39% for human-authored ones. This is a comparison the company made itself, and AI-authored changes are probably smaller by design: the agent refuses to approve large pull requests.
  • The agent may block. MuleSoft took the opposite position because its customers run production workloads on its infrastructure. MuleSoft: trust bar for AI-generated code describes a system that stops a merge when AI provenance is detected and the trust signal falls below threshold. The same acceleration that shortens delivery, the team argues, narrows the window for catching trust problems.
  • The agent is one of three required checks. At Figma, a merge needs human review, static analysis and the agent's pass. Figma: dual-model agents for secure code review runs Claude Code and Codex independently against one shared, precedent-based threat model. Figma reports that the pair catches 75.8% of the vulnerabilities that had slipped past human review and static analysis, at a median cost of $0.50 per review. Precision in the first week was 15%; hand-labelling eight weeks of past pull requests pushed it past 70% within a month.

Two smaller cases show the same pressure outside the largest firms. Indonesia's Halodoc: AI review across a full-stack CI library placed the AI step between reviewer assignment and approval. Its stated aim was that a change should not miss common issues just because the assigned reviewer does not know that stack. Wealthfront: in-house AI code review harness chose to build rather than buy after several years of experiments. The team found the general-purpose harnesses too loosely structured and the dedicated review vendors aimed at the wrong goals.

II. On-call work shifts from investigating to approving

The second pattern is the incident channel. Every case here describes the same change in the engineer's role: the agent does the first investigation, and the human reviews its findings.

Wix: AirBot on-call agent for pipeline failures gives the clearest unit economics in the set. Across 30 Slack channels, Wix reports 675 engineering hours saved per month at about $0.30 per interaction. Of 180 fix pull requests the agent generated in a 30-day window, 28 were merged without changes, about 15%. That rate is a useful anchor: most of the value comes from diagnosis, not from fixes that ship untouched.

Anthropic's CI team reports faster first responses rather than hours saved. Anthropic: Claude Tag as first responder for CI failures says the agent now writes the first situation report in every recent CI incident, a median of 14 minutes after the incident opens. Rollout of a fix sits with a separate agent behind feature flags. Figma applies the same idea to security alerts. Figma: agentic triage of SIEM alerts reports a 71% cut in time to resolution on complex alerts. Figma attributes the 20% drop in on-call pages to retrieval-based deduplication before any agent was added, a reminder that not all of the gain is agentic. Figma also sends bot-authored pull requests to draft through a fixed workflow step, because telling the model to do it proved unreliable.

GitHub shows how long the careful version takes. GitHub: LLM-based detection of leaked passwords reached general availability in October 2024. Getting there took offline evaluation against real customer reports, a second model confirming the first, and a traffic scheduler that GitHub later reused in other products. The result was up to a 94% reduction in false positives.

Salesforce names the skill all of this depends on. Its enablement team's framework for thousands of engineers, Salesforce: agentic engineering proficiency framework, rests on one observation: generating code with AI is easier than judging it, and that is where adoption efforts stall.

III. The vendor is its own first customer in support

Nearly every support deployment with credible numbers comes from a company using its own product internally.

  • ServiceNow: agent resolving Level 1 IT tickets says its agent handles more than 90% of targeted Level 1 help-desk volume end to end. The categories are password resets, access requests and VPN issues, with resolution above 99% in those categories. Forrester called autonomous execution at that level a milestone, after years in which help-desk AI only deflected or routed tickets.
  • Salesforce reports that Salesforce: Agentforce on its own help site handles more than 2.2 million conversations a month in seven languages, against 1.5 million handled by human support engineers. On the employee side, Salesforce: Einstein in Slack for employees reports that one internal concierge app cut average resolution time from 48 hours to 30 minutes.
  • After acquiring Informatica in November 2025, Salesforce had 30 days to stand up a help agent. Informatica: help agent built in 24 days converted about 100,000 documents written for humans into a retrieval base. It now resolves 80% of questions, with 5% escalated. The speed came from reusing an existing ingestion pattern, not from building new infrastructure.
  • GitHub: Qubot internal analytics agent points the same approach at data questions. The team's main finding is that curated context made the agent three times faster at reaching the right answer.

The third-party support cases are thinner and come from the vendors' own case studies. SupportFlow: ECOA AI support rollout reports a 73% cut in first response time within 90 days, plus a detail most accounts leave out: about 4% of users asked for a human even when the AI's answer was right. AssetWorks: four support use cases in 30 days reports a 39% cut in support costs, published by its vendor.

IV. Builders converged on orchestrators and narrow agents

Teams that describe their architecture tell the same story of failure followed by decomposition. Informatica's first design gave many tools to one agent. It picked the wrong tools, ran out of context and produced inconsistent outputs. Informatica: CLAIRE multi-agent system now routes each request through an orchestrator to specialist agents, with fixed tool routing and validation between steps across 50–60 model calls. It reports a 90% task success rate and workflows cut from months to days. Rippling: supervisor with specialist sub-agents shipped the same pattern across its HR, IT and payroll data model in about six months. The supervisor coordinates five to seven sub-agents, and loading skills dynamically cut context bloat by 100–500x.

The gains these teams report come from engineering around the model, not from switching models. Atlassian: Rovo agent harness in code mode moved data iteration out of chained tool calls and into sandboxed code. In internal evaluations this cut latency on complex Jira queries by more than 50%, token use by 55%, and raised accuracy by 30%. Figma's precedent policy, Intercom's per-concern sub-reviewers and GitHub's curated data context are the same move: structure the context, narrow each agent's job, and check the output with something deterministic.

V. Outside engineering, the evidence gets thin

Two other clusters are visible in the set, and both are weaker.

Talent and hiring platforms account for six cases, but most describe what a product can do rather than what happened at a customer. Eightfold: talent intelligence platform is typical: certifications and capabilities, but no operating result. Paradox: Olivia hiring assistant cites a McDonald's case with a tenfold rise in candidate conversion and same-day hiring, but the figures come from a case study distributed by the vendor. The exception is Checkr: fine-tuned small model for background checks. Checkr replaced GPT-4 with a fine-tuned 8-billion-parameter Llama 3 model. Accuracy rose from 88% to 97%, and on messy records from 82% to 85%. Response time fell to half a second, and monthly cost dropped from $7,000–12,000 to about $800. It is the only case in the set where a smaller model was measured beating a frontier one on the job itself.

Regional language models show the same gap between announcement and operation. Sarvam AI: multilingual voice agents in India reports more than 100 million users reached through telecom and government channels. Awarri: Nigeria's government-backed LLM describes a model still being built for five low-resource languages. Most other cases in the cluster are launch notices without usage figures.

Where it is still thin

  • Failures. No case in the set describes a rollback, a withdrawn agent or an incident an agent caused. Intercom's revert rates and Figma's 15% first-week precision are the closest the record gets. A field this confident about measurement should have some published negatives.
  • Engineering outside the US. Of the engineering cases above, only Atlassian (Australia), Wix (Israel) and Halodoc (Indonesia) are based outside the United States. There is no engineering-floor case from China, although nine Chinese Software cases exist in the set.
  • Sales, finance and pricing inside software companies. Apart from Salesforce's lead follow-up, the set has nothing on how software firms use AI in their own revenue, billing or finance work.
  • What "Software" contains. Cases are filed under the supplier's industry, so the set also holds work like Baidu: Apollo Go robotaxi operations and Google DeepMind: Gemini tutoring trial in Sierra Leone. Both are strong deployments, but they belong to transport and education. They are also why the country and company counts above overstate how broad the software industry's own adoption is.

What to watch

  1. Does agent approval spread beyond Intercom? If another large engineering organisation publishes an auto-approval share above 10% by year end, the Intercom model is becoming a norm. If the next adopters follow MuleSoft's blocking model instead, the industry is treating AI authorship as a risk signal, not a productivity signal.
  2. Do the revert-rate comparisons survive size controls? A comparison that matches AI and human changes by size would show whether the roughly tenfold safety gap is real or reflects smaller diffs. If it holds, it becomes the strongest argument for agent approval.
  3. Does the Level 1 help-desk result reach customers? ServiceNow expected general availability in the second half of 2026. Customer-side resolution rates near the internal 90% would show that the internal numbers depend on the product, not on the vendor's unusually clean environment.
  4. Does a second hiring platform publish a measured customer result? If one does before year end, the talent cluster moves from marketing to evidence. If none does, the Checkr pattern — narrow tasks, fine-tuned small models, published costs — is the more credible route for AI in HR software.