The Swedish MASAI randomized trial compared an AI-supported mammography-screening workflow with the existing double-reading workflow. AI helped triage and detect findings; clinicians still took part in interpretation.

An analysis published in 2025 reported 6.4 cancers detected per 1,000 women in the AI group and 5.0 per 1,000 in the control group. That answers a detection question at the time of screening. The study also retained follow-up to observe interval cancers diagnosed between two routine screening rounds. MASAI 2025 study

The primary results published in 2026 reported interval-cancer rates of 1.55 per 1,000 in the AI group and 1.76 per 1,000 in the control group, meeting the trial’s pre-specified criterion for non-inferiority. Although the point estimates differed, the study did not show a statistically significant reduction in interval cancers, nor did it demonstrate lower mortality. MASAI 2026 study abstract

The same trial answers different questions at different stages. How many additional cancers are found at screening, and how many people are diagnosed before the next screening round, require different data and different waiting. For a team preparing to implement medical AI, that waiting is part of validation. Whether the Swedish workflow and findings apply in Taiwan also requires separate assessment.

Pages 145–151 of the 2026 Biotechnology Industry White Paper survey digital-health development, while pages 361–363 address AI governance and applications. Once these efforts reach a hospital, they still have to answer what adoption actually changes.

What a performance table cannot answer

Sensitivity, specificity, and other performance measures answer questions under specified test conditions. Before a hospital decides to adopt a tool, it also needs to know how closely those conditions resemble the intended use.

For image interpretation, a tool may mark suspicious areas or help set the order in which cases are read. Both concern images, but they affect workflow differently. Before validation, the point at which the tool intervenes and the previous practice both need to be explicit.

The comparator also needs to be clear. A study that compares a model with an existing set of labels cannot directly answer whether clinicians complete their work faster after using it. Recording operating time alone is likewise insufficient to determine whether patient health outcomes improved. Every measure has a purpose; conclusions should not reach beyond the evidence.

That is why a claim to “improve healthcare quality” needs further questions. Which part of care improved? Which patients are in scope? How does it differ from the previous practice? A hospital cannot decide how to use a tool from one performance table alone.

Good machine-learning-practice principles jointly issued by the U.S. FDA, Health Canada, and the UK’s MHRA identify the performance of the human-and-AI team as an area of focus. Good Machine Learning Practice principles

When a clinician receives an alert that differs from their own judgment, what information is needed to check it? If every check requires opening another system, that added work belongs in the assessment. Whether the alert is understandable can only be learned by testing it with the people who actually use it.

Implementation may redistribute work as well. A tool that helps clinicians prioritize high-risk cases may change the need for subsequent confirmation and contact. These effects need to be measured in the actual setting before anyone can tell whether time saved is accompanied by other work.

Once frontline staff can report a difficulty, someone must still act on it. Otherwise, a problem that adds a few minutes each day can remain until users work around the system themselves.

Does the original evidence still apply after an update?

Hospital equipment, patient populations, and working practices may change, and software itself may be updated. Validation at implementation therefore needs to connect with subsequent follow-up to see whether the original performance still holds.

FDA’s 2025 draft guidance on lifecycle management and marketing submissions for AI-enabled device software functions discusses arrangements including continuous performance monitoring. As of September 2026, it remains a draft, offering management recommendations and a direction for discussion. FDA 2025 draft guidance

Hospitals and developers need to agree beforehand on which changes will be checked and who decides whether use should pause when a problem is found. Monitoring thresholds must be designed for the product’s use and risk; no single number suits every tool.

Versions also need to be traceable. Under the same product name, the model or its settings may differ before and after an update. If an outcome report does not record version and conditions of use, the next team will struggle to know whether it is using the arrangement that was validated.

Clinical performance, routine operations, and payment are distinct questions at implementation. Evidence that a tool performs well does not mean the hospital has arranged integration and support. Product authorization also does not establish that a particular service is reimbursed by Taiwan’s National Health Insurance.

If cost is limited to a software licence, the extra work at the hospital may disappear from view. Time saved by a tool has to be considered with the time required to use it.

When a pilot ends, the software may remain. Who will maintain it, monitor it, and provide training, and where will the funding come from? Those questions should begin before the pilot.


Sources: 2026 Biotechnology Industry White Paper, Industrial Development Administration, Ministry of Economic Affairs, August 2026, printed pp. 145–151 and 361–363; the original MASAI studies and regulatory documents are linked in the text. The observations on workflow, procurement, and follow-up are the author’s analysis.

Model performance is only one layer of validation
  1. Model performance

    How does it perform on the specified data and task?

  2. Clinical use

    What changes for patients and staff when it enters a hospital workflow?

  3. Ongoing monitoring

    Who tracks and addresses problems when the setting or version changes?

These are distinct validation questions, not evidence that a particular system has passed them. Non-inferiority, superiority and reduced mortality require separate interpretation.