HomeEnterprise AI

Enterprise AI Pilots Are Stress-Testing The Business Behind The Technology

July 31, 2026

Todd Pugh, fractional CIO and former CIO of CSAA Insurance Group, on why enterprise AI failures usually expose broken business logic rather than broken technology.

Enterprise AI Pilots Are Stress-Testing The Business Behind The Technology
Credit: CIOnews

Get the latest from CIOnews.

Enterprise AI, governance, risk, and leadership insights for CIOs, CTOs, CISOs, and technology leaders.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Quote icon
"The agentic agent had executed its guardrail and done exactly what it was told to do as part of the company manual. But what we found was that the company manuals contradicted themselves. That was the aha moment."

Todd Pugh

Principal
@
V2R Advisory

Enterprise AI has left the demo stage, and the carriers, banks, and insurers now trying to run it in production are finding the same thing. The software itself tends to work; the trouble shows up in everything it connects to. When a pilot pushes past a few clicks of automation and into real underwriting or claims work, it starts surfacing problems that were always there: contradictory rules, unclear ownership, processes people had quietly worked around for years. In regulated industries especially, deploying AI turns out to be an audit of how the business actually operates.

Todd Pugh has spent 25-plus years in property and casualty insurance, watching that pattern form. He co-invented a patent for underwriting business owners' policies, and served as EVP and CIO of CSAA Insurance Group, a AAA insurer. He now runs V2R Advisory and sits on the board of Nativeorange, where he helps carriers stand up agentic AI as a platform rather than a pile of point tools. His read on why these programs stall is blunt: teams keep blaming the model when the model is doing exactly what it was told.

"The agentic agent had executed its guardrail and done exactly what it was told to do as part of the company manual. But what we found was that the company manuals contradicted themselves. That was the aha moment." Pugh thinks most of the industry's frustration comes from confusing a model that is wrong with a model faithfully executing broken logic. An agent follows its rules to the letter, so an output that looks wrong is usually a reason to go read the source document it was working from. Treating the failure as a diagnostic is what separates a readiness audit that fixes the underlying problem from a pilot that gets written off as a failed experiment.

  • Mirror on the manual: On a recent engagement, Pugh's customer was convinced the system was hallucinating. "This isn't the outcome or decision that I would have made," was the reaction, he said. "But when we went back and actually looked at the source document, the agentic agent had done exactly what it was told to do. The company manuals contradicted themselves." What looked like a technology defect turned into an opening. "There was a significant opportunity to look more holistically and reevaluate what our underwriting appetite is, what our aperture for what good looks like, and how we want to rate and price this business," he said. "It got back to the fundamentals of the company."

  • Breaking bad behaviors: A second carrier ran into a version of the same thing. For decades it had trained its own market to submit new business a particular way, and the AI pilot forced the question of whether that habit still made sense. "Revisiting that was a really tough conversation, but it was an aha moment as well," Pugh said. "When you bring the two things together, there was an opportunity to really reimagine the approach and get a lot more business value out of the business outcome."

Once the diagnosis is done, the hard part of building something that survives contact with production begins. Pugh's answer is a five-part sequence he wants leaders to work through at the same time, with business, underwriting, technology, risk, and compliance all in the room from day one. He calls the output a North Star. The first two pieces both come down to choosing the right fight: a use case has to clear a bar of real business value, and it has to be tested against the ugly reality of the work rather than a clean happy-day path. "A pilot that saves a few clicks may demonstrate the technology, but that's where a lot of people stop, because they don't see the business value," Pugh said. The back half of the sequence covers who is responsible and what happens after launch: someone has to own the outcome, and the technology has to prove it can run in the carrier's real environment, including missing information and uncertain outputs.

Governance is the piece Pugh is most animated about, having watched it done badly for most of his career, treated as a box checked after the fact. Agentic AI moves too fast for that, and the programs that fail are the ones still treating oversight as a final gate. The carriers getting it right build governance into standardized work so it scales as the AI does. In underwriting, that means the system applies the risk appetite while generating a defensible record for audit and regulatory review, a trust architecture rather than a wall poster.

  • Friction over fantasy: "A strong use case needs to address the material friction: submission intake, document interpretation, appetite screening, validation, prioritization, and so on," Pugh said. "Success must be defined before the pilot begins. Start narrow, but have the end in mind, and make sure it's meaningful enough that you have something to measure." Underwriting submissions, he added, are messy by nature, which is exactly why the messiness belongs in the test. Pull people in early, look at real-world viability, and decide up front where a human has to stay in the loop.

  • Holding the bag: "This gets labeled a technology project, and it's more than that. Someone must own the outcomes," Pugh said. "The pilot should test whether the AI works and whether the company can operate it, which gets to ownership, training, change management, and continuous improvement long after the pilot phase."

  • Compliance as a co-pilot: "Here's what it shouldn't look like, and that's checking boxes," Pugh said. "Think about it as part of your operating model. If you have risk or compliance, guess what, they're in the room. They're part of the North Star, bringing that level of expertise." Done that way, he argued, governance becomes part of how value gets created, adding greater value to the customer, member, or policyholder rather than acting as a tax on velocity.

All of this eventually runs into an architecture problem, because almost no serious AI initiative lives inside a single application anymore. It spans systems, business units, and data sources, and it inherits whatever mess is already there. That arrangement is precisely what makes agents hard to run end to end. Pugh's argument is that an enterprise mindset, treating the agentic layer as a unified platform rather than another point solution, is what finally lets a carrier retire legacy systems and modernize for real.

  • Kicking out the stilts: "When you have a business problem, you go look for a solution, and across the enterprise you end up with very complex, disjointed, often disconnected capabilities," Pugh said. "I've moved away from the word tools. I really think about these agentic AI layers as platforms. There are opportunities to streamline and take out older legacy systems by stitching those layers together. It's reducing total cost of ownership, improving traceability and auditability, and letting underwriters and claims adjusters better understand their business."

How far any of this scales depends on the company, and Pugh draws the contrast himself: the conversation with a tier-one carrier looks different from the one with a small farm mutual in the Northeast. The fundamentals hold, but the way you operationalize them shifts. What stays constant is the discipline: a meaningful starting point, success defined before the pilot, governance and adoption built in on day one, then optimize and repeat. In a regulated business, that operational maturity is the whole game.

"For me it's a back-to-basics moment for the industry," Pugh said. "The winners aren't necessarily the organizations running the most pilots. They're the ones that can turn this into repeatable operating capabilities. And they're thinking North Star, and they're thinking end-to-end."

research report

From the Edge to the Core:
Bringing Agentic AI to the Heart of the Enterprise.