Skip to content
Annuaire
Sections
News

AI agents: the battle shifts from chatbots to task execution

AI agents: the battle shifts from chatbots to task execution
L’essentiel

AI assistants no longer want merely to answer questions: they aim to navigate software and carry out entire assignments. For businesses, the real challenge is deciding which actions to delegate, with what controls and under whose responsibility.

À retenir

AI assistants no longer want merely to answer questions: they aim to navigate software and carry out entire assignments. For businesses, the real challenge is deciding which actions to delegate, with what controls and under whose responsibility.

Reading an email, finding an order, checking an invoice, preparing a refund: for an employee, this is a routine sequence. For artificial intelligence, it represents a change of role. The chatbot offered an answer; the agent claims it can move the case forward. Looking ahead to September 2026, the battle is therefore less about conversation than about the authority to act. That raises a very concrete question for businesses: how far should they let software make decisions and click on their behalf?

From answers to sequences of actions

A conversational assistant primarily produces text. An agent combines a model, tools and a workflow loop: it interprets an objective, chooses an action, observes the result and adjusts its next steps. It can query a knowledge base, call an application programming interface or operate a browser. Autonomy is not necessarily complete: a system that prepares an operation and then waits for human approval already follows this approach.

This shift is grounded in actual announcements. In October 2024, Anthropic introduced a computer-use capability for Claude in public beta: observing the screen, moving the cursor, clicking and entering text. Microsoft has developed its agent offering around Copilot Studio, while Salesforce has launched Agentforce. In January 2025, OpenAI unveiled Operator as a research preview, focused on executing tasks in a browser.

These milestones do not prove that a general-purpose digital employee is ready to start work. They indicate a shift in investment: vendors are seeking to embed their models in workflows. The outlook for September 2026 remains a forward-looking analysis: more agents could become available, without their reliability automatically keeping pace with the proliferation of demonstrations.

The decisive test: a case completed correctly

Consider a refund request. The agent must identify the customer, locate the purchase, consult the applicable terms, check for any exceptions and prepare the transaction. Each step seems simple. Together, they form a chain in which an initial error can contaminate everything that follows. Confusing two orders no longer produces only a wrong answer: it can trigger an unwarranted payment.

Real progress is therefore not measured by the number of clicks performed. It is measured by the proportion of cases handled correctly, the cost of rework and the system’s ability to stop. A successful demonstration on a straightforward workflow says nothing about how the agent will respond to an unreadable invoice, an expired session or two conflicting rules.

This distinction also changes the economics. An agent requires model calls, tools and sometimes a dedicated computing environment. Integration, monitoring and exception handling add to the cost. A task completed quickly is cost-effective only if the time spent checking and correcting it does not consume the gains.

What can reasonably be delegated

Start with reversible tasks

The first candidates are bounded, frequent and verifiable tasks: classifying requests, enriching a customer record using authorized sources, matching documents or drafting a report. The agent produces a result that can be checked without immediately committing the business. Scope matters more than the model’s prestige.

A cautious progression distinguishes three levels:

  • Prepare: gather information and propose an action without changing systems of record.
  • Execute under supervision: make a change after explicit approval of its essential parameters.
  • Execute within defined limits: act independently on authorized operations, with caps, logging and a rollback procedure.

This scale avoids the false choice between automating everything and delegating nothing. An agent can send an acknowledgment on its own but request authorization to amend a contract. It can prepare a supplier order without having the authority to pay for it. Permissions must reflect these differences rather than depend solely on an instruction written in a chat window.

Keep sensitive decisions under human responsibility

Recruitment, credit, healthcare, litigation: when an action directly affects a person’s rights or circumstances, the consequences take on a different scale. Human approval has value only if the reviewer understands the case and has enough time. Mechanically clicking “approve” does not constitute meaningful oversight.

The browser: a powerful shortcut and a weak point

Agents capable of using a graphical interface offer an appealing promise: working with existing software, even without custom integration. This is an advantage in businesses where legacy applications, supplier portals and newer tools coexist. But a moved window, a renamed button or an unexpected dialog box can disrupt the workflow.

Where possible, a structured API connection provides more precise commands and results that are easier to verify. The graphical interface remains useful for filling gaps. One plausible scenario is therefore hybrid agents: robust integrations for critical actions, with the browser handling steps that are less well supported.

The battle also concerns connection standards. Introduced by Anthropic in November 2024, the Model Context Protocol aims to standardize exchanges between AI applications, data and tools. This type of mechanism can facilitate integration. However, it guarantees neither data quality, nor the legitimacy of an action, nor the security of the permissions granted.

Security becomes a question of power

A deceived chatbot can produce nonsense. A deceived agent can act. The risk of malicious prompt injection becomes particularly tangible when it reads web pages, emails or external documents. Content it encounters may try to make it ignore its instructions, disclose information or misuse a tool.

The response cannot rely solely on the model’s willingness to cooperate. Its access rights must be limited, environments isolated, secrets protected and the content it reads kept separate from authorization to act. For a sensitive operation, checks must cover what will actually be executed: the recipient, amount, document sent or data changed.

Activity logs become just as important. The business must be able to reconstruct which tools were called, which approvals were obtained and which changes were made. Above all, it needs a stop button and an incident response procedure. Giving an agent a broadly privileged account to simplify a pilot merely postpones the difficulty until the first serious problem.

The market will be won on exception handling

Business software vendors have an advantage: they know their applications’ objects, rules and permissions. Model providers bring capabilities that span a broader range of tasks. Between the two, businesses will have to weigh tight integration, flexibility and vendor dependence. The best agent will not necessarily be the most spectacular, but the one whose limitations are documented and whose errors can be recovered from.

What next? For September 2026 and beyond, the most credible scenario is gradual delegation rather than the sudden replacement of teams. Businesses would be well advised to choose a narrowly defined process, measure failures as well as successes, and expand permissions only when the results justify it. AI’s next frontier is not simply knowing how to act: it is knowing when to act, under what authority and when to hand control back.

Sur votre appareil

Comprendre cet article

L’analyse utilise l’intelligence locale du navigateur lorsqu’elle existe, sinon un résumé extractif. Le texte n’est envoyé à aucun service extérieur.

Facebook X LinkedIn

Ensuite A lire aussi