The Next Step For AI Agents Is A Clear Path From Human Guidance To Earned Autonomy
Edward Cheng, Senior Advisor for Safe AI Agents at Stanford’s Deliberative Democracy Lab, on how shared rules and decision records let companies judge when an agent can act alone.

Enterprises are onboarding a second workforce. AI agents now pay invoices, file forms, and run transactions that used to wait for a person to approve them, and most companies have deployed them faster than they've built any way to manage them. In a recent EY survey, 91% of senior executives said their organization already runs agentic AI, and 85% admitted that at least some of those systems act without real-time human oversight. The more durable answer works one level up from the individual agent: a governance platform that sets the rules every agent operates under, records the reasoning behind each action it takes, and hands over real autonomy only after the agent has earned it.
Edward Cheng is a Senior Advisor for Safe AI Agents at Stanford University's Deliberative Democracy Lab. He invented the Log-Structured Merge Tree, a data structure that underpins many of the storage engines behind modern big-data and AI systems, and he previously served as Vice President of AI at Oracle and Chief Information Officer at the University of Hong Kong. His current work centers on a question most enterprises are only starting to ask: how to manage autonomous agents once they begin making real decisions.
"Enterprises already have a workforce of human employees, and a human resources department, with the policies that go with it, to manage those people," Cheng said. "Now organizations are bringing in a new class of employees, the AI agents. But they don't have an HR function for that new class, and they don't have policies defined to manage how these agents behave."
That absence is what his work at Stanford sets out to correct. His starting point is the same agreement any organization makes with a new hire or a contractor: define the job, the limits, how much trust the role carries, and how performance gets measured, then build the platform that holds the agent to it. It's the logic of managing a workforce, applied to software that now behaves like one.
Govern from the platform, not the agent: "Most of the time, people focus on the capability of the agent first," Cheng said. "They build an order-to-cash agent, then figure out what guardrails and monitoring to put around it, so governance comes as an afterthought. The disadvantage is that every agent running in your organization ends up with its own separate way of being managed." Bolting controls onto each agent holds up until the fleet grows. A dozen agents becomes a hundred, each with its own logic for what it's allowed to touch and no common place to change the rules. Building the governance as a policy layer first, above every agent, keeps one set of controls in force as the population scales.
Transparency of activity: The platform rests on three ideas, and the first is simple visibility into what an agent is doing. A record of every tool the agent touched and every change it made in the environment gives reviewers a way to trace a decision back to what the agent actually did. "I need to know what work you're doing, what tools you're using, and what changes you've made in the environment," Cheng said. "Transparency is about activities."
Accountability for why: Knowing what an agent did isn't enough to hold it responsible. "In order to hold someone accountable, whether it's an agent or a human being, it's very important that I know not only what you're doing, but why you're doing it," Cheng said. "When you decide to send a check to one of my suppliers, what was the context, and what were the rules you considered at the moment you made that decision?" Capturing that reasoning at the point of action is what keeps an agent's decisions able to stand up to questions later, and what turns a bad call into something a team can learn from instead of just discovering after the fact.
Transparency and accountability both describe work an agent has already done. The harder question is how much of that work it should be allowed to do on its own, and when. Cheng's answer is that an agent shouldn't be handed a high-stakes decision on its start date any more than a new employee would be. It has to earn that authority, and the platform, acting as a control point above the agents, is what measures whether it has.
Trust that's earned: "When an agent faces an ambiguous situation, say a healthcare decision, or a financial decision above a certain threshold, it has to establish its trustworthiness before you automate that decision. The way it does that is by collaborating with a human over a period of time. When you and the human make the same decision, say 99% of the time, that's when you hand it over." The threshold does two things at once. It gives the agent a concrete bar to clear rather than a vague sense of readiness, and it gives the people around it a reason to trust the handoff when it comes. Keeping humans close to the handoff through that period, instead of pulling them out the moment an agent looks capable, is what makes the eventual automation hold up.
How agents earn autonomy: In practice, the trust builds across stages. "In the first stage, the agent just watches the human make the decision and collects that as a training data set," said Cheng. "From that data, we train a machine learning model, and the agent starts predicting the right decision and presenting its recommendation, while the human still decides. When its confidence is high enough, it asks the administrator to authorize it to take over. After that, anything above the confidence threshold it automates, anything below it routes back to a human, and we spot-check the automated decisions to confirm the human still agrees." Cheng points to a transactional agent his team built with PwC as the clearest proof. The agent reads an incoming invoice, checks it against the original order and prior invoices, and clears the match, the same accounts-payable work a team of people would rubber-stamp each month. Treating it like a new hire on probation is what let the automation ship without a leap of faith. "It took about three weeks to automate the order-to-cash agent this way, compared with blindly flipping on the automation and hoping it doesn't make a mistake you catch too late."
For now the approach is early. Stanford is running it internally, PwC is piloting it with a handful of its largest customers, and an open-source version went out to developer communities earlier this month. Cheng's larger argument is that this kind of restraint can't stay optional as agents spread through the enterprise.
"AI safety is the responsibility of every innovator," he said. "These days there's almost no new work being built where someone isn't asking first, 'can I use AI to make this happen?' When we reach for something this powerful, it's on us to think about the safety side, and not repeat the mistake we made with social media 10 or 15 years ago, when we chased the benefit and the excitement without thinking about the responsibility."
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.









