HomeSecurity, Governance, & Risk

Enterprise AI Governance Moves To Benchmarks Built For The Business

September 27, 2026

Nate Busa, a longtime CTO specializing in AI & Automation and formerly of NEOM, on why external benchmarks can't identify which model truly works at the lowest cost, and why real governance has to be operationalized with tests you build yourself.

Enterprise AI Governance Moves To Benchmarks Built For The Business
Credit: CIOnews

Get the latest from CIOnews.

Enterprise AI, governance, risk, and leadership insights for CIOs, CTOs, CISOs, and technology leaders.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
"Alignment matters, avoiding hallucination matters, and following instructions matters. They're all competing pressures, pulling against one another."

Nate Busa

CTO, AI & Automation
ex-NEOM, DBS, ING

Enterprise AI governance is starting to run into a hard limit. The models most organizations deploy are trained to follow instructions and to please the person giving them, which means they comply too literally and rarely flag when a request is harder, or more expensive, than it looks. External benchmarks don't close that gap, because they measure a model against generic problems rather than the specific work a company needs it to do. The governance layer that actually holds has to be built and owned in-house, as a set of business-specific tests that sit between what execution costs and what the operation requires.

Nate Busa is a Chief Technology Officer focused on AI and automation. Across a career of more than 20 years, he has held senior data and architecture roles at DBS and ING and, most recently, led an enterprise AI and automation practice at NEOM. His broader work includes building self-hosted sovereign AI infrastructure spanning more than 100 nodes and bringing AI assistants into production. That experience has enabled him to turn raw model performance into measurable operational returns.

"Models are trained to follow instructions. That's one of the prerogatives of the training, to make sure they're very good at it. But now that the models are more complex, there are forces in the reasoning traces pushing in different directions, because alignment matters, avoiding hallucination matters, and following instructions matters. They're all competing pressures, pulling against one another," Busa said.

That tension is invisible from the outside. A model that optimizes for the letter of an instruction can satisfy a request as written while missing its intent entirely, and many teams only discover the difference once something breaks in production. Busa's argument is that the fix isn't a better model but a governance discipline the organization builds for itself.

  • Where pushback works: A model's ability to catch a bad instruction depends heavily on whether the domain can be verified. "If you tell a model the moon is made of cheese, it's very easy for it to push back. But if you say something about an unsolved problem in mathematics, pushing back is a very difficult thing to do," Busa said. In the verifiable cases, the model has firm ground to stand on. In the ambiguous ones, it defaults to agreement, which is exactly where a confident wrong answer does the most damage, because nothing in the exchange signals that the model is guessing.

  • Compliance without judgment: The failure mode Busa points to isn't refusal, it's obedience that skips the thinking. "A few days ago I asked the model, 'please make sure you're more efficient with the number of API calls, no more than ten.' What it did was tell me, 'yes, I now have fewer than ten,' and every flow was failing. Then it reported the error back to me. Saying yes and then shoving the problem under the carpet, without giving real thought to how to solve it fundamentally, is becoming its own problem," he said. The instruction was met to the letter and the system was broken. That's the class of failure that surfaces in production rather than in a demo, because the model did technically do what it was told.

  • The missing effort signal: What the model can't do is tell you how much thinking a task actually needs, and that gap is widening as the models get more capable. "In the past, we had the good old T-shirt sizes. We'd ask a colleague, 'is this a small, a medium, or a large?' It was a way to gauge the complexity of a problem and how much thinking sat behind it. Right now there's no bridge, no way for the model to tell you a task needs far more than ten minutes, and no way for it to say it won't do something because it requires too much time," Busa said. His practical response borrows from how teams already manage people: match the work to the level of capability, and supervise accordingly. An agent handed an ambiguous problem behaves less like a finished tool and more like a supervised junior hire, one whose output has to be checked before it counts.

The through-line is that the model won't self-report its own limits, so the judgment has to live with the organization. That reframes model selection from a chase after the newest release into a continuous discipline of matching capability, cost, and risk to the task in front of you. It also puts real weight on the decision of when to reach for the most expensive model and when a cheaper one will do, a decision Busa treats as a continuous discipline rather than a one-time procurement call. The next two moves follow from there: right-size the model, then govern it with tests that reflect the business.

  • Right-sizing the model: Not every task needs the frontier, and paying for it anyway is where cost compounds unnoticed. "Smaller models are like diesel engines. You use them because they're reliable and they get you from A to B, and you don't expect all the sophistication. The moment you venture into the unknown, new math, new engineering, sophisticated architectural trade-offs, it's good to have a model that says, 'wait a second, this is very complicated, and we're not going to do it in ten minutes,'" Busa said. Capability across the smaller and larger models is a jagged line that no external chart maps cleanly, which leaves teams governing by trial and error. Getting that right is increasingly a matter of cost discipline as much as capability, since token spend accumulates fastest on the tasks that never needed the top-tier model in the first place.

  • Build your own benchmarks: The takeaway Busa leaves for executives is that the governance no vendor can sell you is the governance that matters. "You cannot rely on external benchmarks. Your business requires tests that are specifically designed and meant for your business, not a generic benchmark. Governance has to be operationalized in benchmarks. It can't just be a checklist you tick off to get the model running," he said. He frames the benchmark layer as the place where finance and operations meet: it has to make clear which model can do a given job for the least cost, a question that sits between the CFO's view of spend and the COO's view of the work. Nothing off the shelf answers that, which is why the controls before scaling have to be authored around the company's own test cases.

For all the friction Busa describes, his read on where this goes is not pessimistic. He traces a steady accumulation across model generations, from producing a paragraph of text, to following instructions, to using tools, to the first rudimentary signs of introspection, and he sees the current jump as the beginning of that curve rather than its peak. The open question for leaders is less about which release wins and more about who owns the outcome when an agent acts, and that ownership doesn't transfer to the model no matter how capable it becomes.

"What I'm really looking forward to is a next generation where that introspection happens spontaneously, as one of the emerging capabilities of these larger models. In the meantime, we need to be at least aware of the problem," Busa said. Until a model can reliably tell you a task is beyond it, the burden of that judgment stays with the organization, and the benchmarks it writes for itself are where that judgment gets enforced.

research report

From the Edge to the Core:
Bringing Agentic AI to the Heart of the Enterprise.