
Every organization I have worked for has delegated decision making down to myself and others. No one walked in with full authority and accountability. We were given a defined responsibility, a set of boundaries, and gradually gained more autonomy and agency over time. We do the same with those that we hire. We accumulate evidence that we can be trusted, just as we do with those that work for us. Over time, that evidence allows us to reduce the oversight.
Human-in-the-loop AI governance seems to follow this same pattern. The agent can make a recommendation to take an action, but a person remains responsible for the decision. The person reviews the output, intervenes when necessary, and retains the formal decision right.
This works as long as the person can meaningfully exercise that right. When the agent can produce decisions or actions faster than the person can reasonably review them, the relationship starts to change. They are not formally removed from the process, but they stop doing deep reviews and instead start doing lighter reviews. Eventually those become spot checks, and ultimately they may stop reviewing the output. This may not happen intentionally or carelessly. We already have more competing for our attention than we have capacity to act upon. If the agent is consistently producing useful results, spending time reviewing every output becomes difficult to justify. Our experience builds confidence which changes how closely we watch. This means that the person may remain accountable while slowly losing the ability to exercise judgment.
This is where our accountability may become performative. We may say that someone remains responsible because they technically have the authority to reject the agent’s output. But the volume makes meaningful review impractical. The human-in-the-loop becomes less useful as a safeguard. The better the AI, the more pronounced this becomes. An agent that produces consistent good results makes reduced scrutiny a reasonable response. The more successful the delegation becomes, the less reason a busy person may have to examine every result.
I was the one responsible for automated back-office jobs in a previous role. In a way, my team was the human-in-the-loop long before that was a term. Most of our jobs ran without issue, and as a result they got very little attention. It was only the ones that had frequent or impactful issues that we really looked into. What I learned doing this work was that some of the required interventions were not actually required. There was one job that handled client notifications. This job would fail after hours, nearly always due to a database lock, and the runbook required us to cancel the alert and re-run the job. This almost always allowed the job to run successfully on the second or third try. One night the job took far more tries than ever to succeed. Around 4 A.M. the job finally finished. I decided that I needed to know why this job required such attention. Only to discover that it didn’t. The job was fulfilling a regulatory requirement that only had to be done once every five business days, we were doing it nightly. I changed how we responded to the job. Mostly by moving the work to business hours.
I think we should consider agents in the same light. If we cannot reasonably observe every decision, the governance has to define how the agent is allowed to act. Scope becomes one of the most important guardrails: what can the agent decide, where can it operate, what consequences are acceptable, and when does a person need to take over.
Those boundaries should reflect the nature of the decision. Low-consequence, reversible decision may use sampling or exception-based reviews. Significant and/or irreversible decisions likely require much stronger interventions. Human-in-the-loop shouldn’t mean examining everything. It should make the interventions intentional.
Even good guardrails do not guarantee a good outcome. They are backstops. An agent can remain within its scope, avoid every tripwire, and still produce an outcome that the organization did not want. That brings us back to accountability.
When we delegate a decision to an agent, have we given the person we say is responsible enough authority and practical ability to actually exercise that decision?