AI now touches three-quarters of enterprise code – New Relic’s Ashan Willy on agent debt and the re-defining of observability
AI now writes the majority of weekly code at most large US enterprises, and the engineer on call when it fails statistically did not write a line of it. The 2026 State of AI Coding Report – a Hanover Research study of 200 US-based technology leaders that New Relic commissioned and published this month – puts the headline number at 67% of leaders saying AI generates 51% to 75% of their weekly code. New Relic’s own observation supports this: GitHub commits jumped three standard deviations above baseline in November 2025, as Copilot, Claude and Cursor became more involved with coding.
Speaking with Ashan Willy, New Relic’s CEO, the interesting question is no longer whether AI is writing the code. It is what observability is supposed to do next, and whether the productivity claims propping up the current adoption cycle survive contact with actual customers. By Willy’s account – they do not.
Willy is prepared to call out productivity inflation in his own industry. Anthropic publicly claims its engineers write 80% of their code with AI and see roughly 10x productivity. Willy’s customer data does not show that. Most people he talks to say it’s about 1.3 right now — a 30% increase because teams are also creating “agent debt.”
Agent debt is the term Brian Emerson, New Relic’s Chief Product Officer, used to anchor the company’s New Relic Now virtual event on June 23. Emerson defined it as “this operational cost that accumulates over time where we’re trying to ship fast, but at some point you run into a challenge of your ability to ship fast if you can’t govern in the right way.”
“The promised productivity gains of AI coding assistants are a mirage if they simply shift the bottleneck from developers writing code to SREs fixing it.”
Willy quotes a recent New Relic study showing the average production outage now costs roughly twice what it did a year ago, moving from about 2 million.
That principle drives the architecture of Preflight, New Relic’s new open-source AI coding observability tool, which reached general availability on June 23 after an announcement at the start of the month. Preflight deploys to a developer’s local machine in around five minutes, watches sub-agent behavior, model usage, token costs and conflicts, and pushes the resulting telemetry into the New Relic platform for team-level analysis.
New Relic treats OpenTelemetry as a first-class citizen for telemetry inside coding agents. Charity Majors’ second edition of Observability Engineering (June 17) formalizes the validation argument: “When agents generate most of your diffs, you can’t validate by reading every line; the proof has to come from production telemetry.”
The Hanover report shows 96% rating observability as essential, with zero rating it unimportant.
Near-term roadmap items include Pathpoint (visualizing non-deterministic AI services in business KPI flows), SRE Agent (causal analysis with visible reasoning chains), Autopilot (domain expert agents in Kubernetes, cloud cost, session replay — late July), and Ground Truth (headless interface for third-party agents to query observability data).
Gartner predicts 90% of enterprise software engineers will be using AI code assistants by 2028, up from less than 14% in early 2024.