Artificial intelligence

AI agent autonomy has to be earned

Gestionnaire confie sa carte de crédit corporative à son employé

Author

Hugues Foltz

Hugues Foltz is Executive Vice-President at Vooban, where he advises Quebec and Canadian business leaders on their applied AI deployments. He speaks regularly on AI adoption in business and on how an organization decides what to delegate to a machine.

Before I hand a new hire my corporate credit card on day one, I'd have quite a few questions to ask.

Recently, an executive told me he'd just plugged a brand-new AI agent straight into his payment systems, to automate his supplier payments. I pointed out that, at the end of the day, that's exactly what he'd just done.

 

I recently shared this conversation on LinkedIn and asked what your biggest questions were about the governance of AI agents. I wrote at the time that it was the start of a line of thinking I wanted to push further. So here is what the post did not say: the three questions to ask, in order, before granting an agent autonomy, the blind spot my own metaphor was hiding, and what I see coming for the organizations that get to work on this now.

An AI agent is a talented new hire. You would not give them the power to authorize payments or sign supplier contracts unsupervised on day one. Yet that is what several companies are doing right now.

I am not saying an agent should never touch your payments. Properly framed, it works. But that is not where you start.

Autonomy Is Calibrated to the Risk of the Decision

What I keep telling executives is that an agent's autonomy has to be earned, and it is calibrated to the risk of the decision it makes.

Many organizations treat this as a single switch. Two camps take shape, and both are wrong.

On one side, people demand that a human approve every move. That destroys the value and recreates the bottleneck you were trying to eliminate.

On the other, people give a blank check. Nobody complains until the first incident.

The mistake both camps share is answering the question globally, when it should be asked agent by agent, and the answer usually sits somewhere in between. An agent that triages minor alerts has nothing in common with one that commits a strategic supplier.

Within the same company, at the same moment, some agents need to be kept on a short leash while others run almost on their own. These are what we call the three supervision regimes, and the regime is chosen one agent at a time.

The Three Questions I Ask

Is the Decision Reversible?

Recommending a stock a little too early is easy to correct, whereas paying an invoice to the wrong supplier is far less so.

Hence my rule: payment systems are the worst place to start. A wire transfer that has gone out is hard to claw back.

How Far Does the Damage Go If It Goes Wrong?

The reach of a decision is measured by the number of operations that depend on it.

Take the agent again. A decision to automate a payment does not just commit an amount: in most enterprise systems, it triggers a bank transfer, updates the accounts payable balance, and adjusts cash flow. One output, three systems hooked into it. In a standard setup, a human intercepts before those three systems move. With an autonomous agent, nothing intercepts anymore.

When a decision reaches that far, you need more control points, not fewer, even when the model is doing its job well.

Do We Have Enough Track Record and Reliable Data to Trust It?

This is the question I see the most executives unable to answer. Most of the time, it was never asked before the agent was plugged in.

In our opening example, the agent has been paying supplier invoices for a few weeks. How many times has it gotten it wrong? By how much? On what type of invoice? If I asked this executive those questions today, he almost certainly would not have any numbers to give me. Not out of negligence, but because no one decided, before go-live, that those numbers needed to exist.

The problem is that without that figure, you are not calibrating anything. You are hoping.

 

Before you calibrate, you first have to know which agents are worth building.

The Blind Spot in My Metaphor

My comparison with a new hire has a blind spot.

A new employee arrives with two things. A mandate, clear before their first day even starts: what they can decide on their own, up to what point, and what must be escalated before they act. And an assigned manager, who answers for their mistakes and can cut off their access on the spot. This executive's agent had neither.

Hence three requirements you have to put in place before any production launch:

  • 1. A written mandate before the agent touches a single system. What it is allowed to decide, up to what point, and what must go through a human. A document that leadership has read and signed, not an instruction buried in the prompt.
  • 2. A named owner, in plain letters, who has the access to act. Not "the AI team." Not "the person who signed the contract." A name, a line in the register, and real technical access.
  • 3. A way to stop the agent without stopping the system around it. What the field calls a kill switch. It gets tested, stopwatch in hand, otherwise it is just a hypothesis.

Done well, agentization makes work more auditable.

These three requirements look heavy until they are in place. Once they are, they slow nothing down: they finally draw the line between what the teams decide and what an agent can settle on its own.

But you still have to hold that line over time. Otherwise, you fall back into the old pattern of Excel spreadsheets stuffed with macros that used to multiply across the company without IT knowing. Except this time, the files make decisions and run on their own.

Without that line, delegation cannot be defended to the executive committee, the audit committee, or in the face of an incident.

What I See Coming

I no longer ask an executive how many agents he has deployed. I ask him to tell me, for each one, what it is allowed to decide. Most cannot.

This know-how is becoming a management skill, even an executive one. In the engagements we see, an agent's error rate is almost never measured. No threshold is set formally. Autonomy does not rise because a bar was cleared, it rises because no one has time to approve anymore and the agent seems to be doing fine. That is escalation by fatigue, not by proof.

And it is already costing. Gartner forecasts that more than 40% of agentic AI projects will be scrapped by the end of 2027, in part because of inadequate risk controls.

My bet is that the organizations that write down their thresholds now, while their fleet of agents is still small, will build a lead the others will take years to close.

So, How Far Do You Let Yours Go?

Your models and your platforms will be replaced within a few years, probably by something better. Your thresholds, your named owners, and your audit logs will remain.

So, before you hand your corporate credit card to your next agent, ask yourself the three questions.

  • 1. Is this decision reversible?

  • 2. How far does the damage go?

  • 3. And do you have a number to justify your trust?

If all three answers reassure you, let it act. If not, stay in the loop.

One last thing. My opening question in the post was about what we are willing to delegate. I will ask it again here, differently: at your organization, which agent is running without anyone knowing who answers for its decisions ?

Setting these thresholds takes work. But we can talk it through.

 

Do you have agents in production and no one has been named owner yet?

To Go Further

Discover how Vooban can transform your projects with innovative technological solutions.

Read more