Blog

New governance framework separates what AI agents can do from what they may do

An academic framework from July 2026 separates the technical capability of AI agents from their permitted autonomy. What does that mean for oversight and control?

· By

Folders with coloured tabs on a wooden desk, ascending from thin to thick, with two open dossiers on the left and a rubber-banded closed dossier and brass key on the right.
A new governance framework separates what an AI agent can technically do from what it may do per workflow.Image: IamVera.ai — original editorial illustration

On 26 July 2026, the paper Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels appeared on arXiv. It introduces a formal distinction between two concepts that are often conflated in practice: what an AI agent can technically do (Autonomous Capability Levels) and what an agent may do in a concrete organisational and risk context (Allowed Autonomy Levels). The authors describe how control, reversibility and accountability must move in step as the permitted autonomy rises.

In practical terms, this means that organisations deploying AI agents can no longer content themselves with the question of how capable a system is. They must set out, per workflow, how far an agent may act independently, who may approve that, and how it can be reconstructed afterwards what happened. The shift is therefore from "AI that talks" to "AI that acts" — and that makes autonomy a design and accountability question.

From talking assistant to agents with formal autonomy levels

The distinction between capability and permission does not stand alone. The survey article Towards Trustworthy Agentic AI: A Comprehensive Survey of Safety and Security Challenges (arXiv, 17 May 2026) summarises several autonomy ladders from the literature. The survey stresses that higher autonomy — especially in systems where multiple agents work together — puts predictability and controllability under pressure and enlarges the attack surface.

An additional picture is offered by the synthesis From LLM Reasoning to Autonomous Agents by AgentMarketCap (5 April 2026). It summarises recent academic work in which agents are broken down into functional layers: perception, planning, action, use of tools and collaboration. It also describes roles running from Operator to Observer, which makes the degree of human oversight explicit at each step. Familiar labels such as assistant, copilot or background agent are thereby rewritten as autonomy tiers with different expectations about reversibility and oversight.

In our assessment, this is the core of the shift: the question is no longer only how powerful an agent is, but at which autonomy level it ran in a specific workflow, and whether that level had been permitted in advance.

Governance architecture around agents that act

Where the academic pieces supply the conceptual framework, practical guides show how organisations make this operational. The AI Agent Data Governance: Enterprise Playbook for 2026 by Promethium (24 April 2026) states that most organisations running agents in production are at an early maturity stage: the agents have considerable operational impact, but governance is still informal. As a foundation, the playbook describes four building blocks: treating agents as their own (non-human) identities with delimited authority, runtime enforcement so that an agent cannot exceed its permitted autonomy, extensive auditing and traceability of data and model lineage.

The practical guide Agentic AI Governance: A Practical 2026 Control Framework by ITECS (16 March 2026) describes how organisations test their deployments against Microsoft's maturity model for agentic AI. According to ITECS, many implementations are still at a low governance level, while documented policy, zoned environments and enforced controls serve as the threshold for responsible scaling. The guide translates autonomy into concrete authorisation tiers: advisory, drafting, bounded action and high-impact action.

As an editorial observation: these sources converge strikingly. The pattern is always identity, enforcement, auditing and lineage — precisely the layers needed to be able to show afterwards who or what did something, with which data and which authority.

Evaluation and verification as autonomy rises

Alongside organisational controls, measurement tooling is emerging. The AgentMarketCap synthesis names the CLASSic framework (Cost, Latency, Accuracy, Security, Stability) for assessing agent stacks along multiple dimensions, plus benchmarks such as GAIA and long-horizon trajectory tests that examine how an agent behaves over longer sequences of steps. This is a shift away from mere chat quality towards reliability and security measures tied to autonomy.

These evaluations do not replace governance; they supplement it. A benchmark says something about behaviour under test conditions, but an organisation must still be able to demonstrate that a specific agent in a specific workflow stayed within its permitted autonomy.

At that point a verification layer such as IamVera.ai fits into the picture — emphatically as a secondary layer, not as the source of the development. Vera is not a chatbot and not its own language model, but a verification layer for professionals working with confidential information. Vera can route a task through selected independent AI models and make verification steps, corrections, mutual disagreements and sources visible. That supports control, but does not guarantee correct outcomes and does not remove the risk of hallucinations; the final judgement remains with the user.

For sensitive documents, the Semantic Privacy Shield can replace values with synthetic, session-only equivalents on EU infrastructure before processing takes place; the architecture is designed to send onward only anonymised content, and when a privacy check fails nothing is sent onward. In the light of the four governance building blocks from the Promethium playbook — identity, enforcement, auditing, lineage — such a verification console mainly gives more insight into the last two: what happened per workflow and which controls applied to it. It is not proof that every autonomous action can be fully reconstructed; it makes inspection possible.

Our conclusion: the common thread in the work of 2026 is that autonomy is becoming a bounded, documented and testable design variable. The question shifts from "how powerful can an agent become" to "how far do we let an agent go, under which conditions, and how do we show that we have maintained that boundary".

Sources: The article draws on arXiv papers about autonomy governance and agentic AI safety and on practical guides from Promethium, ITECS and AgentMarketCap.

← All articles