On-premise medical AI nears cloud performance

September 16, 2026
On-premise medical AI nears cloud performance
AI
News

A medical AI agent running entirely within a healthcare organisation's own infrastructure has achieved diagnostic performance close to a cloud-based model in benchmark testing. Researchers also developed a method to distinguish cases in which the AI appears sufficiently reliable to act autonomously from those that should be escalated to a clinician.

The study, published in Nature Medicine on 15 September, addresses two significant barriers to broader adoption of generative AI in healthcare: maintaining control over sensitive clinical information and determining when an AI-generated decision can actually be trusted. Rather than pursuing complete automation, the researchers propose ‘selective autonomy’, allowing the system to operate independently only when predefined reliability criteria are met.

The findings are promising, but they are not evidence that AI can already diagnose patients independently in routine clinical care. The evaluations were retrospective simulations, making prospective validation in actual clinical environments an essential next step.

Keeping patient data inside the hospital

The researchers developed an autonomous clinical AI agent that can run on-premise, so patient information does not need to be transferred to an external model provider. Such an architecture could give hospitals greater control over privacy, model versions, auditing and the technical environment in which clinical AI operates.

For the primary MIRA-v2 benchmark, which included 551 patient cases across seven diagnostic conditions, four locally deployed open-weight models were compared with GPT-5.2 as a cloud baseline. The best-performing local model, Qwen-3.5, achieved diagnostic accuracy of approximately 90.0 percent, compared with 90.7 percent for the cloud baseline using the same agent architecture and evaluation procedure.

A second benchmark contained 2,400 cases covering four acute abdominal conditions. Here, the system achieved accuracy of 83.8 percent. The researchers also evaluated their approach using VivaBench, consisting of 990 cases across ten medical specialties, to test whether the findings extended beyond the two benchmarks based on MIMIC-IV data.

The relatively small difference between the best local model and the cloud baseline on the primary benchmark is relevant for healthcare organisations considering their AI infrastructure. Local deployment is sometimes assumed to require a substantial sacrifice in model capability. This study suggests that the trade-off may be smaller for certain clinical reasoning tasks, although the result cannot simply be extrapolated to other models, patient populations or applications.

AI must recognise when it is uncertain

Diagnostic accuracy was only one part of the research. The team also evaluated methods for determining how reliable an individual AI-generated diagnosis was likely to be. Behavioural consistency — whether repeated runs of the agent arrived at the same diagnosis — proved more predictive of correctness than several other uncertainty measures examined in the study.

Using a predefined consistency threshold, 49.4 percent of cases could be selected for autonomous handling. Within this selected group, diagnostic accuracy reached 98.9 percent. Cases that failed to reach the threshold could instead be routed to a clinician for further assessment.

This principle could ultimately be more relevant than the headline accuracy figures. Healthcare organisations do not necessarily have to choose between fully autonomous AI and systems that merely provide suggestions to clinicians. Selective autonomy creates a middle ground in which AI can perform defined tasks independently when reliability signals are sufficiently strong, while uncertain cases remain subject to professional judgement.

Clinical practice must be the next test

There are significant limitations. The main benchmarks were derived from MIMIC-IV and therefore represent a particular data environment. The research concentrated on text-based diagnostic reasoning rather than complete multimodal clinical workflows in which an AI agent might directly interpret imaging, laboratory data and other information.

Reliability thresholds were not universal either. Technical settings and the number of repeated runs influenced the results, meaning thresholds would have to be calibrated for individual applications and implementation environments. Bias and performance across different patient groups will also require further assessment.

Most importantly, the evaluations were retrospective. Prospective studies will have to establish how these agents perform when clinicians actually use them and what happens to patient safety, workload, decision-making and clinical outcomes.

Nevertheless, the study points towards an important direction for hospital AI. Keeping clinical data within institutional infrastructure addresses one major implementation barrier, while selective autonomy offers a possible mechanism for determining when human oversight remains necessary. Together, those two elements could prove more important for real-world adoption than simply building another model with a higher benchmark score.

Add ICT&health on Google

Show more content from ICT&health in Google Search.

Add on Google

This topic will also have a prominent place at the ICT&health World Conference 2027. Want to be there and stay ahead of what’s next in healthcare? Reserve your ticket today.