A signed contract, a thousand additional users, rising recurring revenue: on the dashboard, everything looks positive. But every generated response can come with a bill. Every error may require an employee to step in. And behind a promise of automation, there is sometimes a small production team. For an AI start-up, selling more does not necessarily mean earning more. The decisive question becomes: what is left once the intelligence has actually been delivered?
By September 2026, this question could carry greater weight in investors’ and customers’ decisions. This outlook builds on trends already observable in 2024: the spread of open-weight models, price competition among providers and the proliferation of specialized assistants. It is a forward-looking analysis, not a report on financial results that are already known.
Software has a production cost again
Traditional software has never been free to operate. Hosting, support and security consume resources. But serving an additional user can cost relatively little. With generative AI, usage often triggers additional computing costs: document processing, text generation, audio transcription or image analysis. A highly active customer can therefore be less profitable than a light user, even if they pay for the same subscription.
This effect becomes more pronounced when a product chains together several operations. To handle a request, an agent may consult a document repository, call various tools, produce an answer, check it and then start again. A single action visible in the interface can thus conceal several billable calls. The number of user requests is no longer enough to describe the cost of the service.
Competition can nevertheless improve the economics. In 2024, the arrival of OpenAI’s GPT-4o mini and Google’s Gemini 1.5 Flash illustrated the push for lower-cost models for certain uses. Meta’s Llama family, meanwhile, expanded the options for deploying open-weight models. But a lower rate does not guarantee a lower total bill: volumes, long contexts and verification steps can absorb the savings.
The right unit: a successful task
To understand a start-up, it is best to begin with what it actually sells. A usable document? A processed case? A conversation resolved without further intervention? The relevant measure becomes the full cost per successful task, with a definition of success agreed in advance. An instant but unusable response is not a unit of value.
Consider a purely illustrative example. A company charges a customer €1,000 per month. Model calls cost it €200; hosting, storage and the necessary tools, €100; and human corrections directly related to the service, €250. That leaves €450 before other expenses. Counting only the models would show a margin of 80%, compared with a direct contribution margin of 45% under this broader calculation.
This contribution is not net profit. It must still fund research, development, administrative functions and customer acquisition. Accounting classifications can also vary. That is why it is useful to present an accounting gross margin alongside an operational metric reconciled with it, rather than an attractive percentage whose exact scope nobody knows.
What to include in the calculation
- Compute consumed: model calls, unsuccessful attempts, generation, document retrieval and automated checks.
- Infrastructure used: vector databases, storage, monitoring, data transfers and reserved capacity.
- Human production work: validation, correction, moderation and exception handling.
- Customer-specific costs: intensive technical support, uptime commitments and integrations required to deliver the service.
One-off setup costs are best separated from recurring expenses. An expensive deployment may be acceptable if the customer stays for a long time and subsequently becomes profitable. Conversely, a generously priced integration does not prove that the subscription is economically viable.
Humans in the loop—and in the accounts
In healthcare, law or financial operations, human validation may be essential. Elsewhere, it mainly compensates for the product’s limitations. In both cases, it comes at a cost. The problem is not its presence, but its invisibility when the company presents itself as fully automated.
Supervision time per case, escalation rates and the frequency of rework must be tracked. Lower model prices offer little benefit if teams spend more time correcting their responses. Conversely, a more expensive model can improve margins if it sufficiently reduces these interventions.
Beware, too, of founders’ unpaid work. During early pilots, they fix outputs in the evening and personally support every customer. This close involvement helps build the product but distorts its apparent economics. Valuing those hours at a realistic replacement cost helps distinguish temporary learning from a lasting dependence on manual service.
The cloud: supplier, lever and dependency
Credits offered by platforms make the first few months less painful. They do not represent a structural reduction in production costs. A robust dashboard therefore shows margins with and without this support, then models its expiration. The same caution applies to a negotiated discount that requires a substantial spending commitment.
Switching providers is not always straightforward. Formats, tools and model behavior differ. A migration requires quality, security and latency testing. Hosting a model in-house shifts spending to hardware, operations and internal expertise; it does not make it disappear.
Dependency must therefore also be measured in terms of the ability to switch. What share of the service relies on a single provider? How much would a migration cost? What happens to the contribution margin if prices rise or a customer demands a dedicated deployment? These scenarios shed more light on risk than a vague promise of a multi-provider architecture.
Grow margins, not just usage
Technical levers are available: reserve powerful models for difficult cases, shorten instructions, cache certain results and batch non-urgent operations. But every optimization must preserve quality. Reducing the cost of a response while multiplying failures is a false economy.
Pricing matters just as much. An unlimited plan leaves providers exposed to heavy usage; purely usage-based billing worries buyers. A subscription that includes a defined volume, supplemented by transparent overage charges, can distribute risk more effectively. Margins must then be tracked by customer and contract cohort: a reassuring average can conceal persistently loss-making accounts.
What comes next? By September 2026, the strongest start-ups could be those that demonstrate simultaneous improvements in quality, retention and contribution per task. Neither revenue growth nor falling model prices will be enough on their own. The real proof will be a promise fulfilled at a controlled cost, without hidden labor or temporary subsidies.


