Skip to content
Annuaire
Sections
News

Medical AI: Beyond Performance, the Time for Clinical Evidence

Medical AI: Beyond Performance, the Time for Clinical Evidence
L’essentiel

Detecting an abnormality in an image is not enough to improve patient care. To move from laboratory scores to clinical utility, artificial intelligence must demonstrate its benefits in real-world care settings, without shifting risks or increasing workloads.

À retenir

Detecting an abnormality in an image is not enough to improve patient care. To move from laboratory scores to clinical utility, artificial intelligence must demonstrate its benefits in real-world care settings, without shifting risks or increasing workloads.

An alert appears in the medical record: high risk of complications. The algorithm has detected a signal, but what should the nurse, already caring for several patients, do? Call the doctor, take another measurement, wait? The central challenge of medical artificial intelligence lies in this gap between a prediction and a useful decision. Software can pass a statistical test without improving care. Looking ahead to September 2026, the challenge is therefore not simply to build high-performing models: it is to prove that they genuinely help.

The score tells only part of the story

In a study, AI is often evaluated on a collection of images or records for which the final diagnosis is known. It must identify a tumor, predict an infection or classify lesions. Researchers measure its sensitivity, specificity or ability to distinguish patients with the disease from those without it. These metrics are essential. But on their own, they do not describe what will happen in a clinical department.

A database is an orderly world: information is available, categories are defined and results can be verified. Hospitals are less predictable. Test results are missing, schedules are irregular and practices vary. A model may also exploit a shortcut: recognizing the signature of a device or an institution’s practices rather than the disease itself. And if overly similar data appear in both the training and test sets, its performance looks artificially better.

Another pitfall is disease prevalence. Even with good performance metrics, a tool used in a population where the target condition is rare can generate many false alerts. For care teams, this means additional checks; for patients, sometimes anxiety and unnecessary tests. The value of a prediction depends on its context and consequences.

When early results meet real-world practice

Sepsis detection tools illustrate this difficulty. Sepsis is a dangerous response by the body to an infection. An external validation published in 2021 in JAMA Internal Medicine revealed the limitations of a commercial Epic model at the institution studied: missed cases and a substantial alert burden. This finding does not discredit all automated sepsis prediction. It is a reminder that even a widely deployed model must still demonstrate its suitability in local settings.

In breast cancer screening, the Swedish MASAI trial offers a more instructive approach than a simple contest between humans and machines. Its interim results, published in 2023 in The Lancet Oncology, examined AI-assisted mammogram reading within an organized screening program. They suggested a reduction in the reading workload, with at least comparable cancer detection in that analysis. This is encouraging, but does not yet answer every question about long-term benefits.

Detecting more does not always mean providing better care. Some lesions discovered might never have threatened the health of the people concerned: this is the risk of overdiagnosis. Conversely, missing an aggressive lesion can have serious consequences. It is therefore important to examine the types of cancers detected, those diagnosed between screenings, the treatments initiated and, when follow-up allows, the effects on health.

Evaluating a care system, not just an algorithm

The appropriate comparison is generally not between “AI” and “the doctor.” It is between a team equipped with a tool and a team following usual practice. The outcome depends on the interface, when the recommendation appears, training and the time available. An accurate alert displayed after a treatment decision is of little use. A misunderstood recommendation can even cause an error.

Automation bias further complicates the picture: under pressure, a healthcare professional may place too much trust in the software. Conversely, a succession of irrelevant alerts may lead them to ignore a useful one. These reactions are not peripheral details. They are part of how the system actually functions and must be observed during its evaluation.

Steps toward credible evidence

  • Validate elsewhere: test the model at other institutions, with different populations and equipment.
  • Observe without intervening: run it prospectively, without showing its results to care teams, to check its performance on real-world data streams.
  • Compare care pathways: study the tool in actual use, ideally in a randomized trial when appropriate and feasible.
  • Monitor after deployment: look for errors, adverse effects and declines in performance over time.

Randomization can involve patients, departments or institutions. It reduces the risk of attributing to AI an improvement that is actually due to additional staffing or a change in protocol. When randomization is not feasible, other methods are possible, but their limitations must be made explicit. A before-and-after comparison is not always sufficient to establish causality.

Measuring what matters to patients

The primary endpoint must match the promise. If the tool aims to speed up urgent care, the measure should be time to appropriate treatment, not just computing speed. If it automates an administrative task, the time actually saved must be checked, including corrections. If it promises to prevent complications, those complications must be observed, rather than settling for a better risk score.

There is no need to demand a reduction in mortality from every piece of software. A demonstrated decrease in unnecessary tests, better quality of life or faster access to a specialist can all constitute meaningful benefits. But intermediate endpoints must be presented as such. An increase in the number of lesions identified does not automatically demonstrate a reduction in severe forms of the disease.

Averages can also conceal inequalities. Results must be examined across relevant groups: age, sex, comorbidities, geographic origin or equipment characteristics. Sample sizes must support cautious interpretation. A tool that performs satisfactorily across an institution could be less reliable for some of its patients, particularly if they were underrepresented in the training data.

Evidence that must be maintained

Regulatory authorization does not answer every question about purchasing, organization or reimbursement. Each institution must understand the intended use, limitations, validation conditions and full cost: IT integration, training, maintenance and additional checks. Time saved in one department may simply be shifted to another rather than genuinely eliminated.

What next? For September 2026 and beyond, a desirable development would be to make deployments more conditional on verifiable clinical objectives, with stopping rules if the benefits are not confirmed. Every major update should trigger a reassessment proportionate to the risk. The maturity of medical AI will not be measured by the number of spectacular demonstrations, but by its ability to deliver lasting improvements in care, for both care teams and patients.

Sur votre appareil

Comprendre cet article

L’analyse utilise l’intelligence locale du navigateur lorsqu’elle existe, sinon un résumé extractif. Le texte n’est envoyé à aucun service extérieur.

Facebook X LinkedIn

Ensuite A lire aussi