Limited autonomy, ROI still difficult to measure, open architectures, cost control and new governance models: the OECD’s 2026 study shows how the agentic enterprise is actually beginning to take shape.

For the past two years, AI agents have been presented as one of the next major breakthroughs in enterprise artificial intelligence. They are expected to automate entire processes, coordinate multiple applications, develop software, assist scientific research, interact with customers and, ultimately, collaborate with other agents.

Capabilities are progressing rapidly. But a far more important question for leaders remains insufficiently documented: what happens when organizations actually start deploying these systems?

This is precisely the value of the study Agentic AI in organisations: Early insights from practitioner interviews, published in 2026 in the OECD Artificial Intelligence Papers series. Following initial work on the conceptual foundations of agentic AI, the OECD this time shifts the analysis toward practices. The objective is to observe how agentic systems are developed, deployed and governed in real-world contexts (OECD, 2026, pp. 7-8).

The report is based on semi-structured interviews with 25 organizations, including technology developers, user companies, public bodies and research institutions. Anthropic, Google, Microsoft, Salesforce, Deutsche Telekom, Fujitsu, Infosys, NTT, Rakuten, RELX, Sakana AI and the French company Prisme.ai are notably included in the sample (OECD, 2026, pp. 8-9).

The study does not yet demonstrate the widespread economic impact of AI agents. It offers something else, particularly valuable at this stage: a snapshot of early organizational learning. And this snapshot is considerably more nuanced than the promise of an enterprise that has quickly become autonomous.

The agentic enterprise is advancing, but autonomy remains bounded

The first finding deserves leaders’ attention in itself. None of the organizations interviewed report having deployed agentic systems with unrestricted autonomy.

Organizations favor limited autonomy within clearly defined boundaries. High-stakes or difficult-to-reverse actions, such as payments or data deletion, generally remain subject to human validation. The level of intervention varies according to the use case, its risk profile and its environment (OECD, 2026, p. 15).

This finding puts into perspective certain narratives about the imminent advent of enterprises operating through largely autonomous systems. The observed reality is more gradual. Organizations are less focused on eliminating human supervision than on determining where it remains essential.

The report shows, moreover, that this supervision is evolving. The choice is no longer limited to two options: complete autonomy or permanent human validation. Organizations are implementing checkpoints at certain moments in the process and considering gradually increasing autonomy when systems demonstrate their reliability. One organization even uses a five-level framework inspired by the classification of autonomous vehicles, from simple agents to fully autonomous systems within predefined boundaries (OECD, 2026, p. 20).

For leaders, the strategic question therefore changes its formulation. It is no longer about asking how far the technology can become autonomous, but about determining the level of autonomy to assign to each process, based on the value created and the risk accepted.

The best use cases are not necessarily the most spectacular

The second finding is particularly operational. In the short term, the most credible opportunities concern tasks sufficiently structured for an agent to plan and execute them. Their result must be verifiable from reliable data or systems, while the cost of an error must remain limited or the action reversible (OECD, 2026, p. 14).

This framework is important. A company should probably not begin by asking where it could deploy AI agents. It should identify processes in which a capacity for adaptation, judgment or coordination between multiple systems brings additional value.

The OECD provides an essential nuance here: not all tasks that meet these criteria necessarily require an AI agent. Traditional rule-based automation may sometimes suffice. The specific value of agentic AI appears when the task also requires adaptation to changing situations, a certain form of judgment or coordination between different tools and systems (OECD, 2026, p. 14).

Technological sophistication does not in itself constitute value creation.

Agents are beginning to enter operational processes

The study shows that use cases are already numerous. Agents are used for knowledge management, proposal preparation, meeting summaries, data preparation or field intervention organization. Some systems analyze customer requests, search for relevant information in internal documentation and prepare a response subsequently reviewed by a team member (OECD, 2026, p. 11).

Other applications concern regulatory compliance, legal research, document drafting, supply chain and industrial coordination. Software development is among the most frequently cited areas. Agents can participate in code writing, testing, debugging, review or optimization. Cybersecurity is another important field, particularly for vulnerability detection and incident response assistance (OECD, 2026, p. 12).

In telecommunications, agents analyze factors likely to affect demand and can contribute to network capacity planning. In scientific research, ambitions are even greater: hypothesis generation, experiment design, analysis of large amounts of data and, potentially, development of partially autonomous research chains (OECD, 2026, p. 12).

An evolution is taking shape: the AI agent is gradually leaving the conversational interface to penetrate the operational process itself.

“Customer Zero”: the organization becomes its first learning ground

Several B2C-oriented companies interviewed by the OECD have adopted an interesting strategy: testing their agentic systems internally before offering comparable capabilities to their customers. The report refers to “Customer Zero.”

This approach allows validation of performance, identification of unexpected behaviors and gradual building of trust in a more controlled environment. External deployments remain notably more constrained when legal or societal stakes increase (OECD, 2026, p. 14).

This can be seen as an organizational learning strategy: the company uses its own operations as a validation environment before exposing the agent more broadly to its customers.

Scaling no longer depends solely on technical maturity. It also depends on the organization’s ability to learn from its own experiments and transform this learning into operational rules.

The paradox of the agentic enterprise: domain expertise becomes even more strategic

Another finding in the report deserves greater attention. Technical expertise is not sufficient.

The organizations interviewed emphasize the essential role of domain experts in precisely defining tasks, identifying failure scenarios and validating results produced by systems. The OECD even notes that access to this domain knowledge could become a bottleneck as agentic AI develops (OECD, 2026, p. 21).

The paradox of the agentic enterprise may lie here: the more technically powerful agents become, the more strategic domain expertise becomes.

An agent cannot be properly framed, evaluated and governed without those who truly know the process it is supposed to execute. The difficulty may therefore not only be recruiting more AI specialists. It will also be mobilizing team members capable of making sometimes tacit knowledge explicit, identifying exceptions, defining what constitutes an acceptable result and determining situations in which the agent must interrupt its action.

Agentic AI then becomes as much a knowledge organization project as a technology project.

Technical architecture becomes a strategic decision

This is probably one of the most interesting findings in the report. The organizations interviewed do not necessarily seem to want to build their architecture around a single model.

They select or combine models according to tasks, performance, costs, language requirements or their clients’ policies. Some implement mechanisms allowing dynamic switching from one model to another (OECD, 2026, p. 15).

Two expressions used in the study particularly deserve to be retained. The first is “minimum viable intelligence”: using the smallest model capable of correctly accomplishing the task and reserving more powerful models for operations requiring greater reasoning or planning. The second is “strategic optionality”: preserving the ability to change technology when needs, costs or performance evolve (OECD, 2026, p. 15).

The architecture choice then becomes a strategic governance decision. In the agentic enterprise, choosing an architecture capable of changing models could become more strategic than choosing the best model today.

MCP, A2A and the question of technological dependence

This search for optionality also explains the interest in standards and open components.

The report notably mentions the Model Context Protocol (MCP) for connecting agents to tools, as well as agent-to-agent communication protocols such as A2A. The organizations interviewed associate these approaches with interoperability objectives, but also with a strategic concern: avoiding lock-in to a vendor’s ecosystem. Open-weight models are also used to diversify architectures and reduce certain dependencies (OECD, 2026, p. 15).

Competitive advantage may not lie solely in access to the most powerful model. It will also depend on the company’s ability to orchestrate multiple models, agents, tools and data, while maintaining the freedom to evolve its architecture.

Interoperability is then no longer just an IT issue. It becomes a form of strategic insurance against technological dependence.

How to measure an agent’s performance?

This is probably one of the least resolved problems. Evaluation mechanisms for agentic systems remain less mature than those for traditional language models. The OECD notes that no common standard has yet emerged for comprehensively evaluating an agent’s behavior over an extended sequence of actions (OECD, 2026, p. 19).

However, some organizations are experimenting with new metrics: intent resolution, meaning the agent’s ability to correctly understand the request; tool-calling accuracy, its ability to correctly select and parameterize the right tool; task adherence, which measures compliance with the assigned mission; and response completeness, which evaluates the completeness of the result produced (OECD, 2026, p. 21).

But the study also shows the emergence of indicators much closer to economic management: time saved, process throughput, agent utilization, projected cost, carbon footprint or user feedback (OECD, 2026, p. 21).

This evolution is fundamental. The model benchmark is no longer sufficient. An extremely high-performing model can create little value if it is integrated into a poorly designed process. Conversely, an architecture combining several less powerful models can produce more value if it correctly executes a process, at the right cost and with an acceptable level of risk.

For leaders, the ROI of agentic AI should therefore be assessed at the business process level, not solely at the model level.

Autonomy must also have a budget

The study reveals a much more concrete problem: the operational cost of agents.

Unlike a one-time query to a chatbot, an agent can chain multiple model calls, query different tools, operate in loops or be triggered regularly. Its resource consumption can therefore become difficult to anticipate.

The OECD reports that two organizations interviewed did indeed experience cost overruns. The report emphasizes that loops and recurring triggers can cause resource consumption to grow in ways that are difficult to predict (OECD, 2026, p. 20).

This finding opens a dimension of agentic governance that has been little discussed. We talk a lot about functional permissions: what data can the agent access? What tools can it call? What actions can it perform? Economic permissions will likely need to be added.

How long can an agent operate? How many calls can it make? What amount of resources can it consume? What maximum budget can it commit to accomplish a task? Some organizations are beginning precisely to set limits on computation, memory or execution time (OECD, 2026, p. 22).

Autonomy must also have a budget.

“AI inside, rules around it”: governance becomes an architecture

The early governance practices described by the OECD show another evolution. It is no longer just about writing an AI usage policy. Control mechanisms must be integrated into the system architecture.

The organizations interviewed use isolated environments to test agents and limit their rights according to the principles of least privilege and zero trust. They also impose human validation for certain sensitive actions and are beginning to implement agent authentication mechanisms (OECD, 2026, pp. 22-23).

Some also use external layers based on rules. The agent can propose an action, but its final execution is subject to a deterministic system that verifies it complies with planned constraints. The report summarizes this approach with a particularly telling formula: “AI inside, rules around it” (OECD, 2026, p. 22).

The idea is important. It is not about asking the model to fully self-regulate, but about placing predictable, verifiable and auditable mechanisms around it.

The study also mentions registries of authorized agents to limit shadow agents, as well as observability mechanisms capable of tracking actions, decisions and data modifications. Some organizations even use specialized agents to monitor other agents and report abnormal behavior (OECD, 2026, pp. 22-23).

Governance is therefore gradually becoming multilayered.

The human risk: gradually losing what we delegate

In the midst of these technological and organizational questions appears a more subtle issue.

The more powerful agents become, the more users may be tempted to delegate not only execution, but also part of their judgment.

The organizations interviewed mention risks of automation bias, cognitive delegation and cognitive debt. They also emphasize the possibility of gradual erosion of expertise and learning opportunities through experience. One organization even recommends that team members periodically continue to perform certain tasks without agentic assistance in order to maintain their expertise and ability to critically evaluate results produced by systems (OECD, 2026, p. 19).

This observation leads to a question that could become central in the transformation of work:

What skills must the enterprise deliberately keep with humans, even when they become automatable?

The answer does not solely concern human resources. It directly concerns the organization’s resilience. A company that automates a skill to the point of no longer knowing how to exercise it also becomes dependent on the system to which it has delegated it.

A valuable snapshot, but not yet an economic demonstration

However, the study must be interpreted with caution. The OECD explicitly emphasizes its limitations.

The 25 organizations interviewed are already involved in the research, development or deployment of agentic systems. They therefore do not constitute a representative sample of all enterprises. Sectoral and geographical coverage remains limited, maturity levels differ and the speed of technological developments may make some observations rapidly evolving.

Above all, the study does not provide independent quantitative validation of system performance. The results essentially reflect experiences and perceptions reported by the organizations interviewed (OECD, 2026, p. 10).

We must therefore speak of early empirical findings, not definitive proof of the economic performance of agentic AI. This nuance strengthens rather than diminishes the report’s value: it shows how much we are still at the beginning of organizational learning.

From experimentation to execution: the real challenge of 2026

The OECD’s conclusion remains cautious. The report considers that agentic AI is transitioning from an experimental capability to an operational element in enterprises, the public sector and certain consumer-facing environments (OECD, 2026, p. 24).

But organizations are at very different maturity levels. Evaluation and assurance mechanisms remain under construction. Cybersecurity questions persist. Responsibilities between organizations and between agents remain insufficiently defined.

The OECD nevertheless observes the emergence of several common practices: gradual increase in autonomy, human checkpoints for sensitive actions and co-design with domain experts. It particularly emphasizes the need to evolve governance as autonomy increases (OECD, 2026, pp. 23-24).

My reading of this study is therefore as follows: 2025 was largely the year of the agentic promise. 2026 is beginning to become that of organizational learning.

We are starting to move from a discussion of what agents are theoretically capable of doing to a much more demanding question: under what conditions should they actually be allowed to act?

The most advanced organizations do not seem to be seeking maximum autonomy. They are learning to select appropriate processes, frame agents’ actions and maintain checkpoints. They are also seeking to measure process performance, preserve their freedom of technology choice, control costs and retain essential human skills.

This may be where the real shift from experimentation to execution lies.

The agentic enterprise will not necessarily be the one that has deployed the greatest number of agents. It could be the one that knows precisely why it delegates an action to them, how far it allows them to go and how it measures the value created by this delegation.

Artificial intelligence creates a capability. Organizations create the value.