An assistant prepares a refund, opens the billing software and clicks “confirm”. The customer receives their money. Then an employee discovers that the request was fraudulent. Who should answer for the mistake: the model developer, the integrator, the business manager or the company that gave it access to the account? With AI agents, the question is no longer simply whether an answer is correct, but who authorises its consequences. This scenario illustrates the work facing organisations in the run-up to September 2026.
This analysis draws on developments documented up to 2024; the developments envisaged for September 2026 are forward-looking projections. This is an essential distinction in a sector where spectacular demonstrations often precede evidence of reliability.
From advice to execution: the real shift
A conventional chatbot produces text. An agent combines a model with tools: a search engine, messaging system, customer database, computer terminal or payment interface. It can break down a request, select an operation, examine its result and continue. The degree of autonomy varies: some workflows are tightly defined, while others give the model greater latitude.
This trajectory was already apparent in 2023, with features enabling models to call tools, AutoGPT experiments and assistants connected to business applications. In 2024, demonstrations of coding agents extended that promise: moving from “here is how to do it” to “I’ll take care of it”. But a successful demonstration guarantees neither robustness in production nor effective handling of exceptions.
The potential gains are tangible. A procurement department could automate the collection of quotes and prepare an order. An IT support team could diagnose a fault and then restore access. Yet the nature of the risk changes: a poor suggestion can be debated; file deletion, a disclosure or a bank transfer can have consequences before anyone reviews them.
Responsibility is distributed, not dissolved
Saying that “the AI decided” resolves nothing. Technical autonomy alone does not make an agent a legal person to which blame can be transferred. Depending on the country, the contract and the harm caused, several parties may share liability: the company using the agent, a component supplier, the integrator or a service provider responsible for operations.
A Canadian case provided an early warning in 2024. In Moffatt v Air Canada, British Columbia’s Civil Resolution Tribunal held the airline liable for incorrect information its chatbot provided about a bereavement fare. This did not involve an agent performing an autonomous operation. Nevertheless, the case is a reminder that a company cannot simply portray its automated interface as separate from its obligations to customers.
With an agent, the investigation becomes more complex. Did the model misinterpret an instruction? Did the integration give it too much power? Was a business rule missing? Did the employee have enough information to approve the action? Technical cause, legal liability and operational responsibility must be distinguished. They do not necessarily point to the same party.
Permissions become the first safeguard
The trap would be to give an agent the same access rights as the employee it assists. An employee may legitimately access sensitive records, authorise spending and send emails. Their assistant does not need all these capabilities to sort invoices. The principle of least privilege, long established in cybersecurity, becomes a basic requirement for deployment.
In practice, each agent should have a distinct technical identity and limited permissions that are temporary wherever possible. Permission to read does not imply permission to modify; permission to prepare does not imply permission to send; permission to create a payee does not imply permission to pay them. Financial limits, approved recipients and prohibitions must be enforced by the execution system, not merely requested in an instruction to the model.
- Reversible actions: automation is possible within a defined scope, with checks after execution.
- Sensitive actions: explicit approval before sending, publishing or making significant changes.
- Critical actions: segregation of duties, dual approval or exclusion from autonomous execution.
Another threat is the injection of malicious instructions into a document or page the agent accesses. An agent tasked with summarising an email could encounter an instruction within it asking it to export files. This external content must remain data, rather than become a source of authority. Access restrictions and output filtering must limit the damage even if the model is deceived.
Oversight means more than clicking “approve”
Human oversight seems an obvious answer. Yet it can become an organisational fiction. If an employee receives dozens of incomprehensible requests under time pressure, they may end up approving everything. The human then remains in the loop on paper, but no longer exercises meaningful control.
Useful approval presents the planned action, its consequences, the supporting evidence and the relevant uncertainties. For a refund, this means displaying the amount, the recipient, the applicable rule and any anomalies. The supervisor must be able to reject, correct and halt the process without being penalised for slowing execution.
The right question is therefore not “do we have human approval?” but “can a competent person actually prevent harm?” This requires time, training and decision-making authority. Looking ahead to September 2026, this genuine capacity for control can be expected to become a more important selection criterion than the number of tasks promised.
Tracing actions without recording everything
After an incident, a saved conversation is not enough. It must be possible to reconstruct the tools called, the parameters passed, the permissions available, the responses received and the approvals obtained. The model version, agent configuration and applicable rules also matter: the same objective can produce different workflows after an update.
This traceability does not require exposing the model’s supposed exhaustive internal reasoning. Above all, it requires observable evidence of its actions. Logs must be protected against tampering, accessible to authorised personnel and subject to proportionate retention periods. Recording everything indiscriminately would create another risk: accumulating personal data, trade secrets or credentials.
The law sets the framework; organisations must follow through
Adopted in 2024, the EU AI Act provides for phased implementation, with a general application date of 2 August 2026 under its original timetable, subject to exceptions. It does not automatically classify all agents as high-risk systems: their use and the conditions set out in the legislation are decisive. The GDPR, contract law and sector-specific rules also remain relevant. Labelling a system an “agent” does not erase these obligations.
What next? The credible scenario is not a company left at the mercy of all-powerful software, but autonomy that is graduated, tested and revocable. Before expanding an agent’s deployment, executives should be able to answer three questions: what commitments can it make, who can stop it and what evidence will remain tomorrow? Trust will come less from its conversational fluency than from the strength of the limits placed around it.


