Meeting minutes drafted in thirty seconds, a customer response prepared instantly, code produced on demand: AI knows how to make its benefits visible. So do the bills. Between the two, one question defies the demonstrations: how much does the business really gain? Looking ahead to September 2026, the economic challenge is no longer just equipping employees, but verifying that the tools improve financial performance. An hour saved is neither an hour eliminated nor an hour sold. And an inexpensive subscription can conceal a costly undertaking.
Productivity gains are real, but they do not come alone
Available research provides compelling reasons to take an interest in generative AI. A field study published in 2023, involving more than 5,000 support agents, found an average increase of around 14% in the number of issues resolved per hour with generative assistance. The gains were particularly pronounced among less experienced employees. However, this finding relates to a specific environment, not a promise that applies to every occupation.
Another experiment, conducted with Boston Consulting Group consultants and published in 2023, highlighted a crucial boundary: on tasks suited to the model’s capabilities, participants using AI improved in speed and quality. On a task outside that boundary, however, assistance could reduce accuracy. AI therefore does not deliver uniform returns: it redistributes performance according to tasks, people and controls.
For a finance department, the implication is immediate. An average drawn from a study or an internal survey is not enough to build a budget. Businesses must identify the operations involved, measure their frequency and check that the improvement holds up under ordinary conditions: incomplete documents, exceptions, tight deadlines and users less enthusiastic than the pilot’s volunteers.
The real price goes beyond licences
The first cost is easy to identify: a per-user subscription, API usage or reserved computing capacity. The rest is spread across IT, business functions, legal and human resources. Connecting an assistant to internal data requires managing access rights, cleaning up sources and maintaining interfaces. A document retrieved too easily may also be one that should never have been accessible.
Then come training, user support and the time spent checking outputs. A plausible but incorrect answer sometimes means redoing the entire job. For a sensitive task, validation by a qualified professional remains necessary; it must be included in the calculation from the outset, rather than emerging as an unpleasant surprise after launch.
- Launch costs: vendor selection, integration, data preparation, testing and change management.
- Recurring costs: licences, usage, maintenance, support and quality evaluation.
- Oversight costs: review, incident management, security, compliance and error handling.
Comparing subscription prices alone is therefore like comparing two vehicles without considering fuel, maintenance or how they will be used. The right metric is often the fully loaded cost per acceptable outcome: a correctly processed case, a ticket resolved for good or an approved document. It makes visible the rework that sales demonstrations leave out of the picture.
From time saved to value realised
Imagine a department that saves ten minutes on every sales proposal. This is an illustrative calculation, not a market statistic. If salespeople use that time to contact more qualified prospects, a benefit may emerge. But those additional contacts still have to turn into profitable sales. If the real bottleneck lies elsewhere, in product availability or pricing approval, those ten minutes unlock nothing.
Three situations must be distinguished. The first is an actual budget saving: an expense disappears, such as an external service that is no longer needed. The second is additional capacity: the team handles greater volumes without additional hiring. The third is improved quality: shorter turnaround times, better-documented cases and more consistent service. All can have value, but they are not accounted for in the same way.
Multiplying self-reported hours by an hourly wage does not prove a saving. Salaries generally continue to be paid. That calculation estimates theoretical capacity freed up. To claim a return on investment, businesses must demonstrate how that capacity is used: credible cost avoidance, additional margin or reduced losses. They must also avoid counting the same benefit twice, as both a cost reduction and new sales capacity.
Designing measurement that stands up to enthusiasm
Start before installation
A rigorous pilot starts with a snapshot of existing operations. How long does a task take, and how much does that vary? What is its error rate? How many cases come back for correction? Without a baseline, the final measurement mostly captures impressions. Activity logs and quality checks usefully supplement users’ reports, without turning the experiment into continuous monitoring of individuals.
Compare like with like
Where possible, a team using the tools is compared with a control team performing similar tasks. A phased rollout can also facilitate comparison. Seasonality, experience levels and case complexity must be taken into account. Otherwise, a quiet month or an exceptionally motivated team can make a contextual effect look like software performance.
Measure through to the final outcome
Measuring the time taken to draft an email is insufficient if the recipient then has to ask for clarification. In software development, producing more code does not guarantee faster delivery: review, testing and maintenance matter too. In customer service, the number of completed conversations must be assessed alongside reopened cases and satisfaction. The meaningful gain is measured across the entire process.
A dashboard for decisions, not persuasion
The formula remains standard: divide the net benefit attributable to the project by its total cost over a defined period. But the assumptions deserve as much attention as the result. A conservative scenario, a base case and a favourable scenario make uncertainties around adoption, volumes and errors visible. The time needed to recoup the investment usefully complements the percentage return.
Management should also set stopping criteria. A tool that sees little use, cannot meet the expected quality standard or is too costly to supervise does not automatically deserve another phase. Conversely, a modest assistant can be profitable when applied to a repetitive operation. The right unit of decision-making is not “AI in the enterprise”, but a specific use case with an accountable owner, a budget and a verifiable outcome.
What next? For September 2026 and beyond, one plausible scenario is that purchasing decisions will increasingly depend on operational evidence rather than technological prestige. There is no guarantee that this discipline will take hold everywhere. But the best-positioned businesses will probably be those able to distinguish between time freed up, capacity put to use and money actually earned. AI does not need to turn every minute into a euro to be useful; it simply needs to stop presenting those two units as equivalent.


