A client signs up, users flock in, recurring revenue climbs. On the dashboard, the artificial intelligence start-up looks like a success. But behind each response may lie several model calls, a document search, an automated check and sometimes a human correction. Revenue tells the story of traction; margins tell the story of product viability. Looking ahead to September 2026, this distinction could become a decisive screening criterion for entrepreneurs and their backers. Here is how to assess it, drawing on trends documented through June 2024 and their possible implications.
Software that incurs costs with every use
In traditional subscription software, serving an additional user generally costs little once the product has been developed. Generative AI changes that equation: producing a response requires computing resources. Heavy usage, an excellent commercial signal, can therefore undermine the profitability of a poorly calibrated plan. What looks like the best customer sometimes becomes the one who consumes the entire margin.
As early as 2023 and 2024, offerings from OpenAI, Anthropic and Google made this mechanism visible through pricing based notably on the volume of tokens processed. The launch of GPT-4o in May 2024 also illustrated the decline in listed prices between model generations. But a cheaper model does not guarantee a more profitable product: more usage, longer context and additional steps can absorb the savings.
For an assistant summarizing a document, the process may remain short. For a tool tasked with qualifying a lead or preparing a case file, it becomes longer: retrieve the data, interpret the request, call tools, check the response and repeat if necessary. The relevant bill is no longer the cost of a single call, but the cost of the completed work.
Calculating margins at the level of a useful outcome
The first task is to define a meaningful economic unit. Depending on the product, this might be an accepted case file, a resolved conversation, a processed invoice or a validated document. Counting requests alone conceals failures and rework. A system with a low cost per response can become expensive if the user needs several attempts.
The initial calculation is simple: revenue attributable to the unit, minus the direct costs required to deliver it. The next step is to distinguish gross margin, under the accounting rules adopted, from an operating contribution margin that includes expenses that are variable or directly attributable to the service. This second perspective supports decision-making without replacing consistent accounting.
- Compute: model calls, dedicated hosting, GPU capacity and additional processing.
- Data: extraction, indexing, storage, document retrieval and licenses required for the service.
- Human operations: verification, correction, escalations and support directly linked to cases.
- Quality failures: rework, refunds or customer credits attributable to incidents.
This classification must avoid double counting. It must also make temporary integration costs visible without confusing them with ongoing operating costs. A customer may be profitable after deployment but require so much initial work that recouping that investment becomes uncertain.
Humans in the loop—and on the income statement
Imagine a tool that drafts responses for a customer service team. The demonstration is impressive; in production, some responses require review. If that verification is essential to delivering on the commercial promise, its cost belongs in the service’s economics. Permanently classifying it as research or overhead artificially flatters performance.
Companies need to measure the time actually spent, not just the number of escalated cases. A small proportion of difficult cases can monopolize a team. Founders who correct outputs for free in the evening are also providing a resource: to assess scalability, their work must be valued at a realistic replacement cost.
Supervision is not, however, an admission of failure. In sensitive applications, it can provide an assurance customers are willing to pay for. The question is who pays for it and what value it adds. If customers perform the review themselves, the cost disappears from the start-up’s bill, but not from the customer’s return on investment.
Supplier dependence is an economic risk
Using an API makes it possible to launch a product without funding the infrastructure to run its models. In return, part of the business’s economics depends on a third party: pricing, rate limits, availability, contractual terms and changes to model versions. A change can force new testing, disrupt quality or require a migration.
The answer is not necessarily to host your own model. Open-weight models, whose profile the Llama family had helped raise before June 2024, offer additional options. But self-hosting requires expertise, maintenance and capacity that may sometimes be underused. An API with a high unit cost can still be cheaper than a reserved server sitting idle.
The right comparison is total cost at comparable quality. Switching providers requires evaluating results on representative business cases, not just public rankings. And maintaining several options carries its own cost: connectors, testing and monitoring. Portability is insurance, not a free benefit.
A dashboard that exposes costly cases
The overall average can easily provide reassurance. Yet it blends simple and complex customers, short documents and vast archives, controlled usage and unchecked automation. Entrepreneurs need to track costs by customer, feature and task category, then examine the extremes of consumption.
A useful dashboard brings together net revenue, cost per accepted outcome, frequency of rework, minutes of supervision and escalations to a more expensive model. It also tracks how cohorts evolve: do new customers become less costly to serve as experience grows, or do they gradually discover more resource-intensive uses?
These observations should guide pricing. An unlimited plan can work if usage is predictable and can be pooled across customers. Otherwise, quotas, tiers and charges for heavy processing offer better margin protection. Outcome-based pricing can align incentives, provided the outcome and the responsibilities in the event of failure are precisely defined.
Cutting costs without compromising the promise
The technical levers are tangible: reserve powerful models for difficult tasks, shorten context, cache what can be cached and batch non-urgent processing. Sometimes, a deterministic rule or a conventional search can be a better alternative to generation. Every optimization must nevertheless be tested: compute savings canceled out by additional corrections are no savings at all.
What next? For September 2026 and beyond, the plausible scenario is not the end of the race for revenue, but a race constrained by stronger economic evidence. The best-positioned start-ups may be those able to show how much a successfully completed task costs them, how that cost evolves and what happens if their provider changes its terms. The real competitive edge would then no longer be merely a spectacular model, but a useful, repeatable and profitable promise.


