Data Modeling

Mathematical and Computational Data Modeling | Online ISSN 3143-9217
2
Citations
20k
Views
49
Articles
Your new experience awaits. Try the new design now and help us make it even better
Switch to the new experience
Figures and Tables
REVIEWS   (Open Access)

Artificial Intelligence for Quality Control and Error Detection in Diagnostic Laboratories

Saiful Islam 1, Md. Anisur Rahman Anis 2, Md. Abdur Rahim Sarkar 3, Sujibur Rahman 4, Md. Basharuzzaman 5, Faiz un-Nisa 6, Bheesham Kingrani 7, Maryam Zafar 8

+ Author Affiliations

Data Modeling 7 (1) 1-16 https://doi.org/10.25163/data.7110945

Submitted: 11 July 2026 Revised: 13 September 2026  Published: 25 September 2026 


Abstract

Diagnostic error is common, consequential, and, for the most part, difficult to see while it is happening. In the decade since the National Academies of Sciences, Engineering, and Medicine reframed it as a public-health problem, artificial intelligence (AI) has drifted from speculative promise into the working life of diagnostic laboratories, although how far it can be trusted there remains unsettled. This review attempts to take stock. We undertook a structured narrative synthesis, informed by PRISMA-ScR reporting logic, of peer-reviewed literature indexed in PubMed/MEDLINE, Scopus, and Web of Science between January 2019 and July 2026. Twenty-five primary sources were retained, spanning clinical chemistry, microbiology, pathology, cardiology, and adjacent diagnostic fields, and they were synthesised thematically rather than pooled statistically. Reported accuracy or concordance generally fell between 0.78 and 0.99, with the strongest results in narrow, well-defined tasks such as melanoma classification, tuberculosis radiography, and myocardial infarction rule-out. Pathogen-genomics applications reached detection or classification rates of roughly 85% to 97% for organisms including Mycobacterium tuberculosis and Bacillus anthracis. Neurocardiology models achieved AUCs from 0.79 to 0.997, yet performance dropped noticeably once questions became differential or prognostic. Eleven recurring barriers emerged, concentrated around data quality, generalisability, and algorithmic bias, with automation bias and regulatory drift appearing repeatedly as well. Taken together, the evidence suggests that AI can strengthen predictive quality control and error detection, but unevenly. Its safe translation seems to depend less on headline accuracy than on site-specific calibration, explainable outputs, prospective validation, and governance that continues long after deployment.

Keywords: Artificial intelligence; diagnostic error; quality control; clinical laboratory medicine; pathogen genomics; explainable AI

1. Introduction

Medical diagnosis is sometimes imagined as a quiet, back-office affair, the part of care that happens before the real decisions are made. That picture was never quite right, and it is becoming less right every year. Diagnostic work has been moving, gradually at first and then rather quickly, away from observational and largely manual routines toward something more layered, more data-dependent, and, uncomfortably, harder to audit by eye (Hirosawa & Shimizu, 2025). A convenient marker of this shift is the 2015 report of the National Academies of Sciences, Engineering, and Medicine, Improving Diagnosis in Health Care, which defined diagnostic error as the failure to establish an accurate and timely explanation of a patient’s health problem, or to communicate that explanation to the patient (National Academies of Sciences, Engineering, and Medicine [NASEM], 2015, as cited in Shimizu et al., 2025). The definition was not especially elaborate. What mattered, arguably, was what it made visible: errors were not rare or dramatic events but ordinary and recurrent, woven into cognitive shortcuts, fragmented systems, and communication that slips through the cracks (Hall et al., 2020; Singh et al., 2014).

The scale is easy to state and rather harder to absorb. Drawing on three large observational datasets, Singh et al. (2014) estimated that about one in twenty adults in the United States experiences a diagnostic error in outpatient care each year, a proportion that translates into roughly twelve million people. Hall et al. (2020), in their analysis of patient-safety practices for the Agency for Healthcare Research and Quality, likewise counted diagnostic error among the safety concerns that warrant dedicated, systems-level attention rather than reliance on individual vigilance alone.

It is against this background that artificial intelligence (AI) and machine learning (ML) have been put forward, somewhat ambitiously, as the next enabler of diagnostic safety. The optimism is not baseless. AI-driven tools are already being folded into clinical workflows to automate parts of data analysis, to flag anomalies that a tired reviewer might miss, and, at least in principle, to reduce the human error that audits keep uncovering (Al-Antari, 2023; Albahra et al., 2023). The stakes are perhaps highest in in vitro diagnostics (IVD), a field whose results are said to inform somewhere between 60% and 70% of critical clinical decisions (Wurcel et al., 2019). Once that figure settles in, laboratory quality control begins to look less like a compliance exercise and more like a patient-safety intervention in its own right. From interpreting genomic sequences to monitoring sepsis risk on an intensive care unit, AI is increasingly described, sometimes too confidently, as indispensable to both diagnostic accuracy and operational throughput (Bekbolatova et al., 2024; Guzman-Garcia et al., 2025).

Yet the arrival of AI in real clinical environments has not been the clean, linear success story its early advocates may have pictured. Schnetler et al. (2025) put the problem bluntly, describing as a “false hope” the assumption that a single, generalisable model can be moved between hospital sites and simply work, as though data bias were a footnote rather than a determinant of whether an algorithm helps or quietly harms. The field also still lacks standardised methodologies for clinical AI research, which makes studies hard to compare and slows translation into everyday practice (Zhang et al., 2026). Meanwhile, laboratories themselves are consolidating into larger, higher-throughput operations (Vandenberg et al., 2020). Consolidation, somewhat ironically, raises rather than lowers the stakes for quality control. Validation needs to be more rigorous; bias needs active mitigation rather than polite acknowledgement; data-handling practices have to withstand closer scrutiny (Daly et al., 2026); and regulatory instruments such as the European Union’s AI Act must somehow keep pace with technology that does not wait for legislation (Gasser, 2023).

Clinical microbiology shows this tension especially clearly. Pathogen genomics, chiefly whole-genome sequencing (WGS) and metagenomics, offers a resolution for pathogen identification and antimicrobial resistance (AMR) prediction that phenotypic methods struggle to match (Ali & Muhammad, 2023; Gador-Whyte et al., 2026), and AI methods are being applied ever more widely to the diagnosis and prevention of hospital-acquired infection (Baddal et al., 2024). None of this comes cheaply, though. Specialised training, integration with laboratory information systems (LIS), and sustained infrastructure investment are all prerequisites, and they are not evenly distributed across health systems, which is an equity problem in its own right and one that tends to be discussed quietly, if at all. Layered on top of this is generative AI. Large language models (LLMs) such as ChatGPT have found an unexpected foothold in clinical communication, laboratory interpretation, and medical education (Egli, 2023; Niu et al., 2025), and they have even been trialled as expert systems for detecting resistance mechanisms (Giske et al., 2024). Useful, certainly. But there are legitimate concerns about the reliability, and now and then the scientific plausibility, of what these models produce.

Read together, these threads suggest that diagnostic excellence is no longer a matter of better instruments alone. It seems to require a more integrated effort spanning research, education, practice improvement, patient engagement, and policy (Shimizu et al., 2025), together with a sober accounting of weaknesses and threats alongside strengths and opportunities (Sallam et al., 2025). What still appears to be missing, at least in a form that laboratory professionals can act upon, is a synthesis that sets performance evidence beside the structural barriers that decide whether that performance survives contact with routine practice. This review tries to sit at that intersection. Rather than treating AI adoption as an inevitable upgrade, it asks a more grounded question: how might these tools be woven into diagnostic workflows in ways that actually improve patient outcomes, while respecting the ethical and regulatory complexity that any serious innovation in medicine eventually runs into?

Four questions guided the work. First, how do site-specific data distributions and training biases affect the performance and safety of AI-driven prediction models across clinical settings, with sepsis alerting serving as an instructive case? Second, does embedding explainable AI (XAI) into laboratory workflows support frontline diagnostic decision-making better than opaque, “black-box” models do? Third, what barriers limit the equitable global implementation of pathogen genomics, and could decentralised sequencing help curb AMR in low-resource settings? Fourth, could a purpose-built quality-assurance framework, for instance one designed for multiplex quantitative clinical chemistry and laboratory-developed tests, plausibly secure reproducibility and clinical effectiveness at scale?

Building on these questions, the review aims to (i) evaluate the reported effectiveness of current AI applications in improving diagnostic accuracy across laboratory medicine, infectious disease management, and allied specialties; (ii) identify and categorise implementation risks, particularly data privacy, algorithmic bias, and workforce displacement; (iii) outline a roadmap for standardised AI validation that pairs risk-based regulatory approaches with conventional performance metrics; (iv) examine how pathogen genomics is reshaping infection control and decentralised public-health surveillance; and (v) consider how patient engagement and interprofessional education might foster a culture of safety within high-reliability organisations (HROs).

2. Artificial Intelligence in Diagnostic Medicine and Laboratory Practice, Balancing Accuracy, Barriers, and Governance

2.1 Diagnostic Error as a Systems Problem

Diagnosis has long rested on clinical observation, pattern recognition, and accumulated heuristics, and in many respects it still does. What has changed, perhaps faster than many anticipated, is the degree to which these traditional skills are now supplemented by data-intensive approaches oriented toward precision and patient safety (Hirosawa & Shimizu, 2025). The NASEM report is usually credited with much of the recent momentum. In the decade since its publication, international research networks have expanded considerably, and their attention has shifted toward the systemic roots of diagnostic failure: cognitive bias, breakdowns in communication, and the fragmentation of care across settings and providers (Shimizu et al., 2025).

Earlier epidemiological work had already pointed in this direction. The outpatient error estimates reported by Singh et al. (2014) were simply too large to be explained away as isolated individual lapses, and later patient-safety syntheses came to treat diagnosis as a property of the system as much as a skill of the clinician (Hall et al., 2020). For laboratories, this reframing carries a particular edge. Errors can originate before, during, or after the analytical phase, and many of them remain invisible to the clinician who eventually acts on the result. It is partly for this reason, we suspect, that the laboratory has come to be seen as a natural home for predictive quality control and, by extension, for AI.

2.2 From Rule-Based Systems to Large Language Models

AI now sits close to the centre of this conversation. Its intellectual lineage is usually traced to Alan Turing’s 1950 inquiry into whether machines could think and to the 1956 Dartmouth Conference, where the field acquired its name (Niu et al., 2025). Early medical AI systems were largely rule-based, built on explicit if–then logic that was transparent but brittle. Contemporary systems look quite different. They draw on ML, deep convolutional neural networks (CNNs), vision transformers, and LLMs such as GPT-4 and PaLM (Al-Antari, 2023; Niu et al., 2025). Beneath the newer architectures, however, much of what is actually deployed in pathology and laboratory medicine still depends on supervised learning, and its reliability hinges on data-preprocessing decisions that are seldom discussed but are often decisive (Albahra et al., 2023).

Generative models have widened the range of possible applications and, it seems, the range of possible failures too. Egli (2023) asked whether LLMs might represent the next revolution for clinical microbiology; the answer so far looks like a qualified “perhaps.” GPT-4-based agents, for example, have been tested for detecting AMR mechanisms, yet their outputs still appear to require careful expert verification before they could be trusted in routine use (Giske et al., 2024). Broader reviews describe a similar ambivalence: real potential for diagnostics and workflow support, shadowed by unresolved ethical questions and uneven public trust (Bekbolatova et al., 2024; Sallam et al., 2025).

2.3 Laboratory Medicine, In Vitro Diagnostics, and Consolidation

The appeal of AI for diagnostic safety is fairly intuitive. Algorithms can process large volumes of structured laboratory and imaging data, automate complex pattern recognition, and, in principle, reduce (though hardly eliminate) some of the cognitive biases that shape human judgement (Hirosawa & Shimizu, 2025; Niu et al., 2025). Because IVD informs so large a share of clinical decisions (Wurcel et al., 2019), even modest gains in error detection might carry disproportionate downstream effects. Clinical decision support systems built on laboratory data are one obvious route, although Daly et al. (2026) caution that their promise depends heavily on how data quality, missingness, and interoperability are handled in practice. Work in cardiology, orthopaedics, and oncology suggests, too, that such tools can support diagnosis and prevention when they are embedded thoughtfully in care pathways rather than appended to them (Guzman-Garcia et al., 2025).

Structural change in the laboratory sector adds another layer. Vandenberg et al. (2020) describe how the consolidation of clinical microbiology services into large, centralised facilities has coincided with the introduction of transformative technologies. Consolidation may well bring efficiencies of scale. It also concentrates risk, since a miscalibrated algorithm in a high-volume hub reaches far more patients than an isolated error at a single bench. The performance figures reported for these applications are summarised in Table 1 and examined in Section 4.

2.4 Computational Pathology in Cutaneous Melanoma

Cutaneous melanoma is among the most aggressive skin malignancies, and the accuracy and timing of its histopathological evaluation bear directly on survival (Venturi et al., 2025). Conventional light microscopy, however, is complicated by marked biological heterogeneity and by diagnostic criteria that leave room for interpretation. Inter-observer discordance among experts is reported at roughly 10% to 25%, most pronounced in borderline melanocytic lesions, spitzoid neoplasms, and severely dysplastic naevi (Venturi et al., 2025). The spread of whole-slide imaging (WSI) converted glass slides into high-resolution digital assets, a shift that may seem unremarkable but is precisely what made AI-assisted evaluation practical (Serag et al., 2019; Venturi et al., 2025). The resulting computational pipeline, running from digitisation and image quality control through classification, feature extraction, immune mapping, molecular inference, and explainable morphometry to tumour board decision-making, is summarised in Figure 1.

Deep learning architectures, particularly CNNs and multiple-instance learning (MIL) models, have shown diagnostic performance comparable to, and in some studies better than, that of expert dermatopathologists in distinguishing melanoma from benign naevi. Pooled estimates place sensitivity at approximately 89% to 92% and specificity at 90% to 94%, while hybrid models report areas under the receiver operating characteristic curve (AUCs) of 0.96 to 0.98 (Venturi et al., 2025). These are striking figures. Still, they derive largely from curated, retrospective datasets, a caveat that recurs throughout this literature. Classification, moreover, is only part of the picture. Segmentation frameworks such as U-Net and Mask R-CNN are increasingly used to extract the parameters on which staging depends; automated Breslow thickness measurements agree closely with manual ones, and automated detection of mitoses and ulceration offers more objective inputs for American Joint Committee on Cancer (AJCC) staging (Venturi et al., 2025).

AI pipelines appear especially well suited to characterising the tumour microenvironment. Graph-based deep learning can map tumour-infiltrating lymphocytes (TILs) by density and proximity to tumour nests, and the resulting metrics seem to track transcriptomic immune signatures and might help anticipate response to immune checkpoint blockade (Venturi et al., 2025). A related, rather more provocative line of work, often called “molecular histopathology,” attempts to infer genetic alterations directly from haematoxylin and eosin (H&E) slides. The groundwork was laid outside dermatology, when Coudray et al. (2018) showed that deep learning could predict several common mutations from non-small cell lung cancer histopathology images. In melanoma, algorithms predicting BRAF V600 status reach accuracies of roughly 75% to 85%, apparently by detecting subtle morphological correlates such as pleomorphism and architectural disorder (Serag et al., 2019; Venturi et al., 2025). That level is not enough to replace next-generation sequencing (NGS). It might, however, serve as a sensible triage step before formal molecular testing.

Clinical hesitation toward opaque algorithms is understandable, and recent work has responded by favouring interpretable morphometric pipelines. Veronesi et al. (2025) offer a useful example: their nuclei-level ML framework extracted more than six million nuclei from WSIs, evaluated 44 geometric and spatial variables, and reached a diagnostic accuracy of 90.4%. What is arguably more important than the headline accuracy is that the model’s predictions map onto familiar histopathological criteria, such as nuclear pleomorphism, irregular spacing, and anisotropy. Explainable systems of this kind are beginning to find a place in Molecular Tumor Boards (MTBs), where multimodal models can support teams facing grey-zone lesions such as melanocytic tumours of uncertain malignant potential (MELTUMPs) and atypical spitzoid lesions (Venturi et al., 2025). Whether such tools will resolve this ambiguity, or merely describe it with greater precision, is still an open question.

Figure 1. AI-enabled computational pathology workflow for cutaneous melanoma, from slide digitisation to Molecular Tumor Board decision support. The diagram traces eight sequential stages: (1) digitisation of H&E slides into whole-slide images; (2) pre-analytical image quality control and stain normalisation (e.g., HistoQC, PathProfiler, DeepFocus); (3) CNN- and MIL-based diagnostic classification; (4) segmentation of staging features such as Breslow thickness, mitoses, and ulceration; (5) graph-based mapping of tumour-infiltrating lymphocytes; (6) inference of BRAF V600 status from morphology as a triage step before sequencing; (7) explainable nuclei-level morphometry; and (8) multimodal review at the Molecular Tumor Board, where the final decision remains human. Reported performance figures are shown in italics within each stage. The lower panel highlights three threats that condition every stage: batch effects, reliance on retrospective evidence, and automation bias. Synthesised from Browning et al. (2024), Coudray et al. (2018), Serag et al. (2019), Venturi et al. (2025), and Veronesi et al. (2025).

2.5 Neurocardiology and the Heart–Brain Axis

Neurocardiology concerns the bidirectional relationship between the central nervous and cardiovascular systems, and it is commonly organised around two axes. The heart–brain axis covers cardiac sources of cerebrovascular events; the brain–heart axis covers secondary cardiac injury following acute neurological insult (Basem et al., 2025). AI has become an increasingly important tool along both, mainly by accelerating diagnosis and sharpening prevention. The principal applications, with the data modalities and performance figures reported for each, are mapped in Figure 2 and tabulated in Table 4.

Atrial fibrillation (AF) is the leading cause of cardioembolic stroke, and its paroxysmal, often silent course means it is easily missed on a routine electrocardiogram (ECG). Deep neural networks offer a partial workaround. In a widely cited retrospective analysis, Attia et al. (2019) showed that an AI-enabled ECG algorithm identified patients with AF from a single sinus-rhythm ECG with an AUC of approximately 0.87, rising to about 0.90 when all ECGs in the preceding 31-day window were considered. Later work suggests that such models may outperform conventional scores such as CHA₂DS₂-VASc and forecast incident AF over one year with AUCs near 0.85 (Basem et al., 2025). The capability is also moving out of the hospital: FDA-cleared wearable and handheld devices have reported community-screening sensitivities as high as 98.5% and specificities of 91.4% (Basem et al., 2025). This matters for embolic stroke of undetermined source (ESUS), where Choi et al. (2024) found that a deep learning model applied to sinus-rhythm ECGs could identify patients with previously undiagnosed AF; across cryptogenic stroke cohorts, reported AUCs range from roughly 0.81 to 0.88 (Basem et al., 2025).

Along the brain–heart axis, acute events such as ischaemic stroke and aneurysmal subarachnoid haemorrhage can trigger catecholamine surges that lead to myocardial injury, arrhythmia, and Takotsubo syndrome (Basem et al., 2025). Telling Takotsubo syndrome apart from acute myocardial infarction (AMI) is difficult at the bedside. Laumer et al. (2022) trained a temporal CNN on echocardiograms that separated the two with an AUC of 0.79 and an accuracy of 74.8%, outperforming a panel of senior cardiologists (AUC 0.71; accuracy 64.4%). Cardiac magnetic resonance appears to offer a richer substrate still; Cau et al. (2023) reported that ML models combining atrial and ventricular strain with parametric mapping could diagnose Takotsubo cardiomyopathy with considerable accuracy. Not every application is image-based, either. Natural language processing (NLP) of unstructured electronic medical record (EMR) text has classified stroke subtypes with about 80% agreement with expert neurologists (Basem et al., 2025), which is respectable, though it also implies that roughly one case in five would still need human adjudication.

Figure 2. Principal AI applications across the heart–brain and brain–heart axes of neurocardiology. The upper band defines the two axes: cardiac sources of cerebrovascular events (heart → brain, blue) and secondary cardiac injury after acute neurological insult (brain → heart, red). Each row beneath traces an application from its input data (grey), through the AI model and clinical task (blue or red), to the performance reported in the literature (green); the amber box marks the underlying clinical problem of distinguishing Takotsubo syndrome from acute myocardial infarction. Applications include occult atrial fibrillation detection from sinus-rhythm ECGs, wearable community screening, post-ESUS risk stratification, valvular and heart-failure detection, NLP-based stroke subtyping, Takotsubo differentiation, MI rule-out, and mortality forecasting. The footer summarises the performance gradient from acute binary to prognostic tasks. Synthesised from Attia et al. (2019), Basem et al. (2025), Cau et al. (2023), Choi et al. (2024), and Laumer et al. (2022).

2.6 Pathogen Genomics and AI in Clinical Microbiology

In clinical microbiology, the most consequential change of the past decade has arguably been genomic rather than algorithmic, although the two are increasingly entangled. Gador-Whyte et al. (2026) describe how WGS has moved from reference laboratories into routine use, supporting susceptibility prediction for Mycobacterium tuberculosis, outbreak investigation for Clostridioides difficile, and national surveillance of carbapenemase-producing Klebsiella clones. AI and ML add a further layer of interpretation. Ali and Muhammad (2023) argue that ML-based AMR prediction is promising but limited by uneven training data, inconsistent phenotypic reference standards, and the difficulty of interpreting model outputs, challenges that seem to stand between proof of concept and practical implementation.

A systematic review by Baddal et al. (2024) gives a sense of the breadth now covered: deep learning to detect false positives in SARS-CoV-2 RT-PCR fluorescence curves, gradient-boosted models predicting minimum inhibitory concentrations for Salmonella, CNN-based identification of Pseudomonas aeruginosa strains, U-Net detection of Bacillus anthracis in tissue, and even the computational discovery of a structurally novel antibiotic candidate, halicin. These applications are summarised in Table 2. They are not equally mature. Some, such as WGS-based tuberculosis susceptibility testing, have already displaced conventional methods in certain public-health laboratories (Gador-Whyte et al., 2026), whereas others remain experimental. The infrastructure required, as noted earlier, is also unevenly distributed, and consolidation of laboratory services may either ease or worsen that disparity depending on how it is managed (Vandenberg et al., 2020).

2.7 Quality Control, Distribution Shift, and Algorithmic Bias

For all the enthusiasm, moving clinical AI into routine practice runs into substantial technical, legal, and operational barriers (Hirosawa & Shimizu, 2025; Zhang et al., 2026). Perhaps the most persistent is the tacit assumption that a single, uncalibrated model will perform uniformly across institutions. Multicentre evaluations repeatedly suggest otherwise. In a retrospective cohort of 969,292 admissions across nine facilities, Schnetler et al. (2025) found that baseline sepsis-prediction models produced excessive false alerts at sites not represented in training, a pattern they attributed to unmitigated site-specific distribution shifts and to location-dependent variation in care.

Digital pathology faces a comparable problem in the form of batch effects. Differences between scanners, variation in H&E staining intensity, and out-of-focus artefacts can lower CNN classification accuracy by 10 to 23 percentage points (Venturi et al., 2025). In response, automated quality-control tools such as PathProfiler, HistoQC, and DeepFocus are increasingly built into laboratory workflows to flag unusable regions and normalise staining before diagnostic inference; Browning et al. (2024), for instance, evaluated PathProfiler within a working diagnostic pathology service. The problem is not confined to images. Vocal biomarkers appear sensitive to ambient acoustic noise in clinical environments, which argues for standardised signal-quality preprocessing before such tools are relied upon (Augusto, 2025). The broader point, made forcefully by Zhang et al. (2026), is that without standardised evaluation methods these failure modes are difficult even to detect, let alone compare across studies.

2.8 Human Factors, Automation Bias, and Ethics

Introducing AI into clinical workflows also brings human-factor dynamics that are not always easy to predict. Survey data suggest that 32% of frontline laboratory staff worry about job displacement through automation, compared with 25% of laboratory managers (Niu et al., 2025). The gap is modest, but it hints that anxiety is distributed unevenly across professional hierarchies. A more immediate clinical risk is automation bias, the tendency to accept algorithmic output uncritically even when it conflicts with clinical judgement or the underlying data (Niu et al., 2025; Venturi et al., 2025). The danger seems greatest in exactly those grey-zone cases where AI is most tempting to consult.

Explainable components, such as feature-attribution heatmaps and calibrated confidence scores, are often proposed as a remedy, on the grounds that they support meaningful human-in-the-loop oversight without undermining clinical autonomy (Niu et al., 2025). Serag et al. (2019) frame the goal differently, as “intelligence augmentation” rather than automation, with repetitive and visually taxing work offloaded to the model while judgement stays with the clinician. Ethical concerns run alongside these practical ones. Privacy and consent obligations constrain how laboratory data can be shared and reused (Daly et al., 2026), and public attitudes toward AI in healthcare remain mixed (Bekbolatova et al., 2024). Whether explanations actually change clinician behaviour, rather than simply reassuring users, has so far received less empirical scrutiny than one might hope.

2.9 Regulatory Governance and Standardised Evaluation

Regulators have begun to respond with formal frameworks. The European Union’s AI Act, published in the Official Journal in July 2024 and in force from August 2024, adopts a risk-based approach under which most AI diagnostic tools fall into the high-risk tier, bringing obligations relating to data governance, bias mitigation, human oversight, transparency, and post-market surveillance (Gasser, 2023; Niu et al., 2025). In the United States, the Food and Drug Administration (FDA) oversees such systems through its Software as a Medical Device (SaMD) action plan, which emphasises Total Product Lifecycle (TPLC) oversight for algorithms that continue to learn after deployment (Augusto, 2025; Niu et al., 2025).

Alongside these regulatory instruments, structured tools have emerged to audit the accuracy and completeness of AI-generated health information before clinical use, most notably the METRICS checklist and the CLEAR tool (Niu et al., 2025). How these layers fit together, linking the barriers described above to laboratory-level mitigations, regulatory instruments, and a longer-term improvement agenda, is sketched in Figure 3. The structural barriers themselves are catalogued in Table 3.

Figure 3. Integrated framework linking structural barriers, laboratory-level mitigations, and governance instruments to a five-domain roadmap for validated AI practice. The left column lists recurring structural barriers identified in this review (see Table 3); the centre column pairs each with a mitigation that can be applied at laboratory level, such as federated learning, site-specific calibration, automated quality control, explainable outputs, interoperability standards, and real-world validation. The right column summarises the governance and evaluation instruments that frame these mitigations, including the EU AI Act, the FDA Software as a Medical Device Total Product Lifecycle approach, and the METRICS and CLEAR evaluation tools. The lower panel maps the five domains of the diagnostic-excellence agenda (research, education, practice improvement, patient engagement, and policy) to a 2035 horizon, and the dashed loop indicates that post-market surveillance feeds back into recalibration. Synthesised from Augusto (2025), Browning et al. (2024), Daly et al. (2026), Gasser (2023), Niu et al. (2025), Schnetler et al. (2025), Shimizu et al. (2025), and Zhang et al. (2026).

2.10 Gaps in the Existing Literature

Several gaps stand out from this body of work. Much of the performance evidence is retrospective and drawn from curated datasets, so it says relatively little about behaviour under real-world distribution shift (Schnetler et al., 2025; Venturi et al., 2025). Studies of explainability tend to describe what an explanation looks like rather than whether it improves decisions. Pathogen-genomics research is concentrated in well-resourced settings, leaving the question of decentralised implementation largely open (Ali & Muhammad, 2023; Gador-Whyte et al., 2026). And there is, as yet, little that brings accuracy data, barriers, and governance together within a single frame that a laboratory director could use. The present review was designed with these gaps in mind.

3. Methods

3.1 Study Design and Reporting Framework

This study synthesises existing evidence; it does not generate new primary data. We adopted a structured narrative-review design whose search, selection, and reporting steps were informed by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) (Tricco et al., 2018). A narrative rather than a systematic meta-analytic design was chosen deliberately. The questions set out in the Introduction cut across specialties, AI architectures, and outcome types, and we judged that forcing them into a single quantitative frame would obscure more than it revealed. The PRISMA-ScR logic was retained, however, so that the search could, in principle, be reproduced by another reviewer working from the same protocol. The overall workflow, from identification to synthesis, is shown in Figure 4.

Figure 4. Workflow of literature identification, selection, data extraction, and thematic synthesis used in this review. The flow diagram, informed by PRISMA-ScR reporting logic (Tricco et al., 2018), shows the sequential stages of the review. Records were identified from PubMed/MEDLINE (primary), Scopus, and Web of Science for January 2019 to July 2026 and supplemented by backward citation tracking; pooled records were de-duplicated, screened by title and abstract, and assessed in full text against eligibility criteria (a) to (c), with reasons for exclusion shown at right. Twenty-five primary sources were included and extracted into a standardised template capturing specialty, condition, architecture, modality, metrics, comparator, and barriers. Thematic synthesis then organised the evidence into four domains, corresponding to Tables 1 to 4, read across four cross-cutting threads: bias, transparency, regulation, and reproducibility.

3.2 Information Sources

PubMed/MEDLINE served as the primary bibliographic database. Scopus and Web of Science were searched as secondary sources to reduce gaps arising from differences in journal indexing, particularly for engineering and computing venues that publish clinical AI work. These database searches were supplemented by backward citation tracking, in which the reference lists of key reviews were screened by hand for additional eligible studies (Basem et al., 2025; Hirosawa & Shimizu, 2025; Zhang et al., 2026).

3.3 Search Strategy

The PubMed search combined Medical Subject Headings (MeSH) with free-text terms in three conceptual blocks, covering the technology, the diagnostic or quality-control context, and the evaluative focus. The string was structured as follows:

("artificial intelligence"[MeSH] OR "machine learning"[MeSH] OR "deep learning") AND ("diagnostic errors"[MeSH] OR "quality control"[MeSH] OR "laboratory medicine" OR "clinical microbiology" OR "pathogen genomics") AND ("validation" OR "predictive" OR "decision support")

Filters restricted results to articles published between 1 January 2019 and 31 July 2026, written in English, and concerning human diagnostic applications. Equivalent strings were adapted to the syntax of Scopus (TITLE-ABS-KEY fields) and Web of Science (Topic field). The 2019 lower bound was chosen because it broadly coincides with the period in which deep learning tools began to be evaluated in clinical laboratory settings at scale; earlier landmark studies were consulted for background where necessary (e.g., Coudray et al., 2018) but were not counted among the included primary sources.

3.4 Eligibility Criteria

Records were included if they (a) reported empirical performance data, namely sensitivity, specificity, accuracy, AUC, or concordance, for an AI or ML model applied to a diagnostic, quality-control, or laboratory-safety task; (b) were published in a peer-reviewed journal or peer-reviewed proceedings, or as an examined thesis; and (c) addressed at least one of clinical chemistry, microbiology, pathology, cardiology, or pharmacovigilance. Comparative benchmarks from closely allied diagnostic specialties (for example, radiology and gastroenterology) were retained where they were reported within an eligible source, because they help to situate laboratory findings. Editorials and commentaries without original data, non-English texts, and preprints lacking peer review were excluded. A small number of theses were retained because of their methodological detail on signal quality and regulatory compliance (Augusto, 2025).

3.5 Study Selection

Retrieved records were pooled and de-duplicated. Selection then proceeded in two sequential stages: screening of titles and abstracts, followed by full-text assessment against the eligibility criteria. Both stages were carried out by the review team, and disagreements were resolved through discussion until consensus was reached. Because the design was narrative rather than systematic and meta-analytic, formal inter-rater reliability statistics (such as Cohen’s kappa) were not calculated; we recognise this as a limitation and return to it in Section 5.8. Twenty-five primary sources met the final criteria and form the evidentiary basis for Tables 1 through 4 (Figure 4).

3.6 Data Extraction

Data were extracted into a standardised spreadsheet template that had been agreed before extraction began. For each included source we recorded the clinical specialty, the target condition, the AI architecture, the input data modality, the reported performance metrics (accuracy, sensitivity, specificity, AUC, or concordance), the comparator where one was reported, and any structural or ethical barrier the authors discussed. For pathogen-genomics studies, the laboratory setting and the principal clinical stakeholder were also captured. Values were transcribed as reported in the source; no metrics were recalculated, and ranges were preserved where authors gave them. The template allowed cross-tabulation into the four summary tables presented in Section 4.

3.7 Data Synthesis

Pooling effect sizes statistically would, in our view, have been indefensible given the heterogeneity of tasks, architectures, reference standards, and outcome metrics. Findings were therefore synthesised thematically. Studies were first grouped by evidence domain (diagnostic accuracy, pathogen genomics, structural barriers, and neurocardiology) and then read across for recurring conceptual threads, namely bias, transparency, regulation, and reproducibility. This approach mirrors that used in comparable narrative syntheses of AI in laboratory medicine and clinical practice (Niu et al., 2025; Sallam et al., 2025). Where a performance figure came from a retrospective or single-site dataset and the source made this clear, we noted it during synthesis so that the strength of the claim could be tempered accordingly.

3.8 Quality Appraisal and Reproducibility

No formal risk-of-bias scoring, such as QUADAS-2, was performed. We state this explicitly rather than gloss over it, because it limits the confidence that can be attached to any single performance estimate. To support reproducibility, the databases, full search string, date limits, filters, and eligibility criteria are reported above in enough detail that another reviewer could rerun the PubMed query and expect a substantially overlapping result set. All sources are cited in APA (7th edition) format, with digital object identifiers where available.

3.9 Ethical Considerations

This review analysed published literature only. It involved no human participants, no patient records, and no identifiable data, so ethical approval and informed consent were not required. All extracted data are presented in Tables 1 to 4.

4. Synthesis of Findings: Performance, Pathogen Genomics, Barriers, and the Heart–Brain Axis

4.1 Overview of the Evidence Base

The twenty-five included sources clustered into four evidence domains (Figure 4): diagnostic accuracy across clinical specialties (Table 1), AI-supported pathogen genomics (Table 2), structural barriers to clinical translation (Table 3), and neurocardiology (Table 4). Most performance data came from retrospective analyses or from reviews that pooled such analyses, and prospective, multi-site evaluations were comparatively scarce. That imbalance shapes nearly everything that follows, and it is worth keeping in mind when reading the figures below.

4.2 Diagnostic Accuracy Across Clinical Specialties

Taken together, the studies suggest that AI performance is not uniformly excellent but is consistently competitive, and in a handful of domains it appears superior to established practice. Reported accuracy or concordance ranged from about 0.78 in fine-grained pathology tasks, such as mitosis detection, to 0.99 in breast-cancer metastasis screening and EEG-based depression classification (Table 1). CNN-family architectures dominated image-based tasks, including dermatological lesion classification, histopathological mitosis detection, and chest radiograph interpretation for tuberculosis, with sensitivity and specificity frequently exceeding 0.90 (Hirosawa & Shimizu, 2025; Venturi et al., 2025). Outside imaging, a random-forest approach to ECG noise filtering reported an accuracy of 0.97, with sensitivity of 0.99 and specificity of 0.96 (Guzman-Garcia et al., 2025).

What stands out, perhaps, is less the ceiling of performance than its variation with task complexity. Coarse-grained classification, such as melanoma versus benign lesion, achieved consistently higher concordance than fine-grained staging tasks such as Breslow thickness estimation or mitosis counting, where subtle morphological cues seem to resist even well-trained networks (Table 1; Figure 1) (Venturi et al., 2025). The distinction matters clinically. It is staging, not simple classification, that often determines treatment intensity.

Table 1. Reported diagnostic performance of artificial intelligence models across clinical specialties. Summary of ten AI applications drawn from the included sources, spanning dermatology, pathology, radiology, cardiology, psychiatry, oncology, gastroenterology, critical care, and microbiology. For each, the table lists the target condition, model architecture, diagnostic task, and the accuracy or concordance, sensitivity, and specificity reported by the original authors; values are transcribed as published, not recalculated, and ranges are retained where given. Comparisons against human readers or unassisted practice are shown in parentheses where available. Figures derive largely from retrospective datasets and should be read as indicative of capability rather than of performance in routine deployment. Abbreviations: AMR, antimicrobial resistance; AUC/AUROC, area under the receiver operating characteristic curve; CNN, convolutional neural network; DL, deep learning; EEG, electroencephalography; ECG, electrocardiogram; ESBL, extended-spectrum beta-lactamase; MIL, multiple-instance learning; ML, machine learning; N/R, not reported; TB, tuberculosis.

Medical specialty

Target condition

AI method / model

Diagnostic task

Accuracy / concordance

Sensitivity

Specificity

Reference

Dermatology

Cutaneous melanoma

CNN, MIL

Lesion classification

0.85–0.96

0.89–0.92

0.90–0.94

Venturi et al. (2025)

Pathology

Melanoma staging

U-Net, Mask R-CNN

Mitosis detection

0.78–0.88

0.82–0.87

0.77–0.86

Venturi et al. (2025)

Pathology

Breast cancer

CNN (CAMELYON16)

Metastasis detection

0.99 (AUC)

0.91 (vs. 0.73 human)

N/R

Hirosawa & Shimizu (2025)

Radiology

Pulmonary TB

DL algorithm

Chest X-ray analysis

0.97

0.92

0.98

Hirosawa & Shimizu (2025)

Cardiology

ECG arrhythmia

Random forest

Noise filtering

0.9726

0.9853

0.9625

Guzman-Garcia et al. (2025)

Psychiatry

Depression

EEG-based CNN

Diagnostic screening

0.99

N/R

N/R

Hirosawa & Shimizu (2025)

Oncology

Cancer treatment

IBM Watson

Treatment-planning concordance

0.96

N/R

N/R

Hirosawa & Shimizu (2025)

Gastroenterology

Colon adenoma

Real-time AI

Adenoma detection rate

37% (AI) vs. 25% (control)

N/R

N/R

Hirosawa & Shimizu (2025)

Critical care

Hospital sepsis

Baseline ML

Prediction alerting

AUROC 0.8496

0.70 (target)

N/R

Schnetler et al. (2025)

Microbiology

AMR detection

Customised GPT-4

Beta-lactamase flagging

0.819 (concordance)

100% (carbapenemase)

69.2% (ESBL)

Giske et al. (2024)

4.3 Comparators and the Sensitivity–Specificity Trade-off

Some of the most informative results were those measured against practice rather than against a fixed ground truth. Real-time AI assistance during colonoscopy raised adenoma detection rates from roughly 25% to 37% compared with unassisted procedures, and in metastasis detection the reported sensitivity of an AI system was 0.91 against 0.73 for human readers (Table 1) (Hirosawa & Shimizu, 2025). Comparisons of this kind arguably say more about clinical usefulness than any isolated AUC.

In microbiology, GPT-4-based expert-system agents identified carbapenemase activity with 100% sensitivity, but specificity for extended-spectrum beta-lactamases (ESBLs) lagged at 69.2%, with overall concordance of 0.819 (Table 1) (Giske et al., 2024). A similar shape appeared in critical care, where baseline sepsis models achieved an AUROC of about 0.85 yet generated excessive false alerts once deployed beyond their training sites (Schnetler et al., 2025). The pattern recurs often enough to be worth naming: AI tends to be very good at not missing the dangerous case and considerably less good at avoiding false alarms. In a busy laboratory, the latter can erode trust surprisingly quickly.

4.4 Pathogen Genomics and AI in Infectious Disease Management

The second cluster of evidence concerns pathogen genomics, where WGS is steadily displacing slower phenotypic methods. For M. tuberculosis, WGS-based susceptibility prediction has, in some public-health laboratories, effectively replaced routine phenotypic drug-susceptibility testing, and genomic approaches have also tailored therapy for severe Staphylococcus aureus infection and uncovered unrecognised C. difficile transmission (Table 2) (Gador-Whyte et al., 2026). Where quantitative performance was reported, detection or classification rates clustered between roughly 85% and 97%, with U-Net detection of B. anthracis in tissue slides reaching approximately 97% and CNN-based identification of P. aeruginosa strains reaching 90.7% (Baddal et al., 2024).

One finding reads almost like science fiction, though it is not. A message-passing neural network screening chemical libraries identified halicin, a structurally novel antibiotic candidate, entirely in silico (Table 2) (Baddal et al., 2024). It is a reminder that AI’s contribution to infectious disease may extend beyond faster diagnosis into drug discovery itself, at least experimentally. At the more modest end of the spectrum, AI-assisted digital reading of stool thick smears for neglected tropical diseases was reported to reduce manual read-out errors in resource-limited settings (Wurcel et al., 2019). It is a less glamorous application, but possibly the one with the largest equity dividend.

Table 2. Applications of AI and genomic technologies in pathogen identification, resistance prediction, and infection control. Ten pathogen- or disease-specific use cases showing how whole-genome sequencing, deep learning, and machine-learning classifiers are being applied across public-health, hospital, diagnostic, research, and resource-limited laboratory settings. For each application the table identifies the technology used, the clinical purpose, the laboratory setting, the principal finding reported, the type of input data, and the stakeholder group most directly served. Maturity varies considerably: some applications (e.g., WGS-based tuberculosis susceptibility testing) have entered routine practice, whereas others (e.g., computational antibiotic discovery) remain experimental. Abbreviations: AMR, antimicrobial resistance; CNN, convolutional neural network; CP-CRE, carbapenemase-producing carbapenem-resistant Enterobacterales; ID, infectious diseases; M&E, monitoring and evaluation; MIC, minimum inhibitory concentration; NN, neural network; qPCR, quantitative polymerase chain reaction; WGS, whole-genome sequencing; WHO, World Health Organization.

Pathogen

Method / technology

Clinical use case

Laboratory setting

Key clinical finding

Data input type

Primary stakeholder

Reference

M. tuberculosis

WGS

Susceptibility prediction

Public-health laboratory

Replaced routine phenotypic testing

Genomic sequence

Pulmonologists

Gador-Whyte et al. (2026)

S. aureus

Bacterial genomics

Treatment failure

Hospital laboratory

Tailored therapy for severe infections

Genomic + clinical

ID physicians

Gador-Whyte et al. (2026)

C. difficile

WGS

Outbreak detection

Tertiary hospital

Identified unrecognised transmission

Genomic sequence

Infection control

Gador-Whyte et al. (2026)

Klebsiella spp.

WGS

AMR surveillance

National network

Tracked spread of CP-CRE clones

Multicentre sequence

Public health

Gador-Whyte et al. (2026)

SARS-CoV-2

Deep learning

RT-PCR interpretation

Diagnostic laboratory

Detected false positives in qPCR

Fluorescence curves

Laboratory technicians

Baddal et al. (2024)

Salmonella

XGBoost ML

MIC prediction

Clinical laboratory

Predicted MICs for 15 drugs

Genomic regions

Microbiologists

Baddal et al. (2024)

P. aeruginosa

CNN

Strain identification

Research / clinical

90.7% classification accuracy

Colony images

Epidemiologists

Baddal et al. (2024)

B. anthracis

U-Net

Pathogen detection

Histopathology

≈97% detection in tissue slides

Digital images

Biosecurity units

Baddal et al. (2024)

E. coli

Message-passing NN

Antibiotic discovery

Computational laboratory

Discovered halicin (novel antibiotic candidate)

Chemical libraries

R&D scientists

Baddal et al. (2024)

Neglected tropical diseases

AI-assisted digital pathology

Epidemiological M&E

Resource-limited setting

Reduced manual read-out errors

Stool thick smears

WHO programmes

Wurcel et al. (2019)

4.5 Structural Barriers to Clinical Translation

If the accuracy figures look encouraging, the barriers documented across the literature go some way toward explaining why routine adoption still lags behind bench-level promise. Thematic synthesis identified eleven recurring barrier themes: data quality, generalisability, algorithmic bias, transparency, regulatory friction, environmental interference, interoperability, reproducibility, education gaps, clinical-utility mismatch, and privacy. For presentation, these were consolidated into ten categories in Table 3, with generalisability captured within the data-quality and bias rows. The barriers, their technical causes, and the mitigations proposed for each are mapped in Figure 3. Data quality and generalisability were cited most often, echoing the argument that a model trained at one site rarely transfers cleanly to another without recalibration (Schnetler et al., 2025; Zhang et al., 2026).

Two barriers merit particular emphasis, because they concern the people around the algorithm more than the algorithm itself. Automation bias, in which clinicians defer too readily to AI output instead of interrogating it, recurred across nearly every domain reviewed (Table 3) (Niu et al., 2025; Zhang et al., 2026). The proposed corrective is conceptually simple, if operationally hard: treating AI as intelligence augmentation rather than as an autonomous decision-maker (Serag et al., 2019). Regulatory friction was the second recurring theme. Frameworks such as the EU AI Act and the FDA’s TPLC approach are maturing (Augusto, 2025; Gasser, 2023), yet prolonged approval timelines and drift in adaptive algorithms remain unresolved tensions rather than mere bureaucratic inconvenience. Privacy constraints, including HIPAA and GDPR compliance, further limit the data sharing on which multi-site validation depends (Daly et al., 2026).

Table 3. Structural barriers to the clinical implementation of AI in diagnostic laboratories. Ten categories of technical, human, and ethical–legal obstacles identified through thematic synthesis, consolidating eleven recurring barrier themes (generalisability is captured within the data-quality and algorithmic-bias rows). For each category, the table specifies the barrier, its effect on laboratory or clinical practice, its underlying technical cause, a mitigation proposed in the literature, the associated human factor, and the principal ethical or legal concern. The table is intended to be read alongside Figure 3, which links these barriers to laboratory-level mitigations and governance instruments. Abbreviations: AUC, area under the curve; EU AI Act, European Union Artificial Intelligence Act; FDA, US Food and Drug Administration; GDPR, General Data Protection Regulation; HIPAA, Health Insurance Portability and Accountability Act; HL7-FHIR, Health Level Seven Fast Healthcare Interoperability Resources; LIS, laboratory information system; QCVS, quality control of vocal signals; XAI, explainable AI.

Challenge category

Specific barrier

Impact on practice

Technical cause

Proposed mitigation

Human factor

Ethical / legal concern

Reference

Data quality

Small datasets

Poor generalisability

Overfitting on noise

Federated learning

Trust erosion

Scientific integrity

Zhang et al. (2026)

Algorithmic bias

Training gap

Disparity in accuracy across sites and groups

Non-representative cohorts

Site-specific calibration

Automation bias

Health inequity

Zhang et al. (2026)

Transparency

“Black-box” logic

Clinician hesitancy

Opaque deep learning

Explainable AI (XAI)

Cognitive load

Accountability

Niu et al. (2025)

Regulatory

EU AI Act / FDA requirements

Prolonged timelines

Adaptive-algorithm drift

Change-control plans

Risk perception

Medico-legal liability

Augusto (2025)

Environment

Ambient noise

Diagnostic error

Acoustic interference

Signal quality control (QCVS)

User literacy

Patient safety

Augusto (2025)

Interoperability

Legacy systems

Workflow disruption

Fragmented infrastructure

HL7-FHIR / LIS integration

Staff resistance

Data security

Niu et al. (2025)

Reproducibility

Batch effects

Inflated metrics

Staining and scanner variance

Stain normalisation

Expert discordance

Peer-review gap

Venturi et al. (2025)

Education

AI literacy gap

Fear of job displacement

Rapid technological change

Simulation training

Burnout

Workforce stability

Basem et al. (2025)

Clinical utility

Endpoint mismatch

Tool disuse

AUC vs. patient outcome

Real-world validation

Practice detachment

Resource waste

Zhang et al. (2026)

Privacy

Sensitive data exposure

Limits on data sharing

HIPAA / GDPR compliance

Synthetic data

Patient anxiety

Informed consent

Daly et al. (2026)

4.6 AI Across the Heart–Brain and Brain–Heart Axes

The fourth cluster illustrates how performance metrics can look almost too good on paper unless one knows where to look for the caveats. Across ten applications spanning the heart–brain and brain–heart axes, reported AUCs ranged from 0.79 to 0.997 (Table 4; Figure 2) (Basem et al., 2025). The highest figure, an AUC of 0.997 for STEMI-Net’s rule-out of myocardial infarction at triage, is impressive. It should, however, be read alongside more modest results elsewhere in the same table. Mortality forecasting with gradient-boosting ensembles reached only 0.7962, and Takotsubo differentiation from AMI, although better than a panel of cardiologists, reached an AUC of about 0.79 (Basem et al., 2025; Laumer et al., 2022).

Wearable-enabled detection deserves a note of its own because it moves AI out of the hospital entirely. Consumer and medical-grade devices running FDA-cleared algorithms can flag paroxysmal AF in the community with a negative predictive value of around 97.5% (Table 4) (Basem et al., 2025). This is one of the more concrete, already-deployed examples of AI acting as preventive infrastructure rather than as a laboratory curiosity, and it connects directly to the problem of occult AF after ESUS (Choi et al., 2024).

Table 4. AI applications linking cardiac and neurological outcomes across the heart–brain and brain–heart axes. Ten AI models applied to neurocardiology, grouped by clinical axis: heart–brain (cardiac sources of cerebrovascular events), brain–heart (cardiac injury after acute neurological insult), and related cardiac, valvular, acute, diagnostic, and prognostic tasks. For each model the table lists the target disorder, algorithm type, input data modality, primary clinical goal, comparison with clinicians or standard methods, and the reported AUC, C-statistic, or other performance measure. Performance is highest for acute binary decisions (e.g., MI rule-out) and lower for differential and prognostic questions (e.g., Takotsubo differentiation, mortality forecasting). Abbreviations: AF, atrial fibrillation; AUC, area under the curve; CHA₂DS₂-VASc, stroke-risk score; CNN, convolutional neural network; CTA, computed tomography angiography; DNN, deep neural network; ECG, electrocardiogram; EMR, electronic medical record; GBM, gradient-boosting machine; HFpEF, heart failure with preserved ejection fraction; MI, myocardial infarction; NLP, natural language processing; NPV, negative predictive value; SVM, support vector machine.

Clinical axis

Disorder / event

AI algorithm type

Data modality (input)

Primary clinical goal

vs. clinician / standard

AUC / performance

Reference

Heart–brain

Atrial fibrillation

Deep learning (DNN)

12-lead ECG (sinus rhythm)

Future-event prediction

Superior to CHA₂DS₂-VASc

0.85–0.90

Basem et al. (2025)

Heart–brain

Ischaemic stroke

NLP / random forest

EMR and text notes

Subtype classification

80% expert agreement

0.80+ (C-statistic)

Basem et al. (2025)

Brain–heart

Takotsubo syndrome

Temporal CNN

Echocardiogram

Differentiate from MI

Superior accuracy

0.79 (vs. 0.71 expert)

Basem et al. (2025)

Cardiac

Heart failure (HFpEF)

HeartBEiT (vision transformer)

ECG waveform

Ejection-fraction detection

Higher than standard CNN

Superior sensitivity

Basem et al. (2025)

Valvular

Mitral stenosis

Gradient boosting

Clinical metrics

Stroke-risk prediction

Superior to logistic regression

0.8037

Basem et al. (2025)

Valvular

Cardiac murmurs

Deep learning

Digital stethoscope

Abnormality detection

FDA-cleared (Eko)

High specificity

Basem et al. (2025)

Acute event

Myocardial infarction

CNN (STEMI-Net)

12-lead ECG

Rule-out at triage

Comparable to experts

0.997

Basem et al. (2025)

Diagnostic

Coronary artery disease

Support vector machine (SVM)

Coronary CTA

Stenosis diagnosis

Superior to three readers

0.94

Basem et al. (2025)

Heart–brain

Cryptogenic stroke

Mobile ECG AI

Wearable device

Paroxysmal AF detection

Increased community detection

High NPV (97.5%)

Basem et al. (2025)

Predictive

Mortality risk

GBM ensemble

Multimodal metrics

Outcome forecasting

Outperforms statistical models

0.7962

Basem et al. (2025)

 

4.7 Cross-Cutting Synthesis

Read side by side, the four bodies of evidence share a consistent shape. AI tends to perform best where the diagnostic question is narrow and acute (is this a STEMI, is this isolate carbapenemase-producing, is this lesion melanoma) and worst, or at least most variably, where the task demands staging, prognosis, or generalisation to unfamiliar populations and equipment (Tables 1 and 4; Figures 1 and 2). That gradient, more than any single accuracy figure, seems to us the more useful takeaway for laboratories deciding where to invest first.

5. From Bench-Level Accuracy Toward Validated Laboratory Practice

5.1 Principal Findings

This review set out to examine whether AI can support predictive quality control and diagnostic error detection in clinical and microbiology laboratories, and what stands in the way. The short answer is a cautious yes, with conditions. Across twenty-five sources, AI systems reached accuracy or concordance figures that were, for the most part, competitive with expert performance (Table 1), and genomic and AI tools were already reshaping parts of infectious disease practice (Table 2). At the same time, eleven recurring barriers (Table 3) and a clear performance gradient between acute binary tasks and prognostic or differential ones (Table 4) suggest that the gap between demonstrated capability and dependable practice is still wide. These findings broadly agree with earlier narrative reviews (Al-Antari, 2023; Hirosawa & Shimizu, 2025), although our synthesis places more weight on the laboratory-specific conditions under which performance holds or fails.

5.2 Accuracy Is Necessary but Not Sufficient

A recurring temptation in this literature is to treat a high AUC as evidence of clinical value. The studies reviewed here suggest that this is, at best, an incomplete inference. The colonoscopy data, in which adenoma detection rose from about 25% to 37% (Hirosawa & Shimizu, 2025), are persuasive precisely because they measure a change in practice rather than agreement with a reference label. By contrast, the comparison between gradient-boosting models and logistic regression in mitral valve disease (AUC 0.80–0.83 versus 0.41) probably says as much about the weakness of the comparator, which performed below chance, as about the strength of the model (Basem et al., 2025). Zhang et al. (2026) describe this as an endpoint mismatch, in which metrics optimised during development do not map neatly onto patient outcomes (Table 3). Laboratories evaluating AI tools might reasonably ask vendors not only how accurate a model is, but what happened to patients, workflows, and error rates when it was switched on.

5.3 Generalisability and the Case for Site-Specific Calibration

The first research question asked how site-specific data distributions and training biases affect AI performance and safety. The clearest answer comes from Schnetler et al. (2025), whose nine-site sepsis cohort showed that a model performing acceptably in aggregate can produce excessive false alerts at unfamiliar sites. Similar degradation appears in digital pathology, where scanner and staining differences can reduce accuracy by up to 23 percentage points (Venturi et al., 2025), and in vocal biomarkers exposed to ambient noise (Augusto, 2025). The common thread is that performance is a property of a model in a setting, not of the model alone.

The practical implication, we think, is that local validation and calibration should be treated as part of implementation rather than as an optional extra. Automated pre-analytical quality control, of the kind evaluated by Browning et al. (2024), offers one way to catch distribution shift before it reaches a diagnosis (Figure 1). Federated learning and larger, more diverse training cohorts may reduce the problem at source (Table 3) (Zhang et al., 2026), though they are unlikely to remove it. For consolidated, high-throughput laboratories (Vandenberg et al., 2020), the case for continuous monitoring is arguably stronger still, because the consequences of silent drift scale with volume.

5.4 Explainability and Keeping the Human in the Loop

The second question concerned whether XAI improves frontline decision-making compared with black-box models. Here the evidence is suggestive rather than conclusive. Interpretable morphometric pipelines, such as the nuclei-level framework of Veronesi et al. (2025), produce outputs that pathologists can map onto familiar criteria, and this seems likely to aid adoption in multidisciplinary settings such as MTBs (Figure 1) (Venturi et al., 2025). The Takotsubo study by Laumer et al. (2022) shows that an algorithm can outperform experienced clinicians on a hard differential question, which makes it all the more important that clinicians understand when it might be wrong.

What remains largely untested is whether explanations change behaviour in the way intended. An explanation that reassures without informing could, in principle, deepen automation bias rather than reduce it (Niu et al., 2025). We would therefore argue for evaluating XAI components by their effect on decisions, for instance on the rate at which clinicians appropriately override incorrect outputs, rather than by their visual plausibility. Framing AI as intelligence augmentation (Serag et al., 2019), with calibrated confidence scores and explicit escalation pathways, seems a sensible interim position while that evidence accumulates.

5.5 Pathogen Genomics, Equity, and Decentralised Surveillance

The third question asked what limits equitable implementation of pathogen genomics, and whether decentralised sequencing might help curb AMR in low-resource settings. The included literature identifies the barriers fairly clearly: specialised training, LIS integration, sustained investment, and data standards that are not yet harmonised (Ali & Muhammad, 2023; Gador-Whyte et al., 2026). It says much less about decentralised deployment. The AI-assisted reading of stool smears for neglected tropical diseases (Wurcel et al., 2019) and the national Klebsiella surveillance network (Gador-Whyte et al., 2026) hint at what distributed models might look like (Table 2), but neither directly tests whether decentralised sequencing reduces resistance. We would characterise this as a plausible hypothesis that the current evidence neither confirms nor refutes. Generative tools may eventually lower the interpretive burden of genomic reports (Egli, 2023), although the specificity problems seen with GPT-4 agents (Giske et al., 2024) suggest caution.

5.6 Quality Assurance for Laboratory-Developed and Multiplex Assays

The fourth question, concerning a purpose-built quality-assurance framework for multiplex quantitative clinical chemistry and laboratory-developed tests, is the one the current literature addresses least directly. None of the included sources evaluated such a framework for proteomic assays specifically, and it would be misleading to claim otherwise. What the evidence does offer is a set of transferable principles: automated pre-analytical quality checks (Browning et al., 2024), standardised signal-quality preprocessing (Augusto, 2025), lifecycle oversight with predetermined change-control plans (Augusto, 2025; Niu et al., 2025), and reporting standards that allow cross-study comparison (Zhang et al., 2026). A framework built on these principles seems feasible in outline (Figure 3). Whether it would secure reproducibility at scale remains, for now, an empirical question that warrants dedicated study.

5.7 Toward a Roadmap for Validated Practice

Realising the potential described in this review will depend on closing the gap between retrospective proof-of-concept studies and prospective evidence of real-world benefit (Hirosawa & Shimizu, 2025; Zhang et al., 2026). The agenda set out by Shimizu et al. (2025), spanning five interrelated domains with a horizon extending to 2035, offers a helpful frame, and Figure 3 attempts to connect it to the barriers and mitigations identified here. In research, prospective validation across diverse, multi-institutional cohorts is needed to expose and correct batch effects and algorithmic bias (Schnetler et al., 2025). In education, broader AI literacy among laboratory and clinical staff may temper both automation bias and anxieties about displacement (Niu et al., 2025). In practice improvement, automated quality control and explainable pipelines need to be built into workflows rather than layered on top of them (Browning et al., 2024).

The remaining two domains connect most directly to the culture of safety that high-reliability organisations aim to sustain. Patient engagement calls for transparent, patient-centred communication of AI-derived findings, including their uncertainty (Bekbolatova et al., 2024; Shimizu et al., 2025). Interprofessional education, linking laboratory scientists, clinicians, data scientists, and quality managers, seems likely to be what turns these principles into habits (Hall et al., 2020). And in policy, adaptive regulatory pathways such as the EU AI Act and the FDA’s SaMD framework will need to balance the pace of innovation against the demands of patient safety, with post-market surveillance treated as a continuing obligation rather than a formality (Gasser, 2023; Niu et al., 2025). None of these domains is likely to be sufficient on its own.

5.8  Limitations of this study 

The review has several strengths. It brings performance evidence, barriers, and governance into a single frame; it spans laboratory medicine, microbiology, pathology, and cardiology rather than a single specialty; and it reports its search strategy in enough detail to be rerun. Its limitations are, however, real. The narrative design and the absence of formal risk-of-bias appraisal mean that individual performance figures should be read as indicative rather than definitive. Inter-rater agreement during screening was not quantified. Several included sources were themselves reviews, so some primary data were seen at one remove, and heterogeneity of metrics prevented quantitative pooling. Pharmacovigilance, although within scope, ended up thinly represented in the final evidence base. Restricting the search to English-language publications may also have under-represented work from regions where implementation barriers are most acute, which is a somewhat uncomfortable irony given the equity concerns raised above.

5.9 Implications and Future Research

For laboratory directors, the findings point toward a practical sequence: begin with narrow, high-stakes binary tasks where AI performance is strongest; insist on local validation before go-live; build automated quality control into the pre-analytical phase; and monitor real-world performance continuously after deployment. For researchers, the most pressing needs appear to be prospective multi-site trials with patient-relevant endpoints, behavioural studies of whether explanations improve decisions, implementation studies of decentralised genomics in low-resource settings, and validation frameworks for multiplex and laboratory-developed assays. For regulators and professional bodies, harmonised reporting standards would make the next review of this kind considerably easier to write (Zhang et al., 2026).

6. Conclusion

AI appears to have earned a place in predictive quality control and diagnostic error detection, not as a distant possibility but as a demonstrated, if uneven, capability across melanoma pathology, pathogen genomics, and neurocardiology. The evidence nonetheless resists a simple story told through accuracy figures alone. Performance falls as tasks move from acute binary decisions toward staging and prognosis; data quality, generalisability, and bias persist as barriers; and automation bias remains a human-factors risk rather than a footnote. Much of this still rests on retrospective data. For laboratories considering adoption, the more defensible path probably lies in combining site-specific validation, explainable outputs, automated quality control, and regulatory alignment, and in continuing to monitor performance after deployment. Treated less as a replacement for expert judgement and more as a closely supervised partner, AI might yet help laboratories reduce diagnostic error. That outcome, however, will have to be earned rather than assumed.

References


Al-Antari, M. A. (2023). Artificial intelligence for medical diagnostics—Existing and future AI technology. Diagnostics, 13(4), Article 688. https://doi.org/10.3390/diagnostics13040688

Albahra, S., Gorbett, T., Robertson, S., D’Aleo, G., Kumar, S. V. S., Ockunzzi, S., Lallo, D., Hu, B., & Rashidi, H. H. (2023). Artificial intelligence and machine learning overview in pathology & laboratory medicine: A general review of data preprocessing and basic supervised concepts. Seminars in Diagnostic Pathology, 40(2), 71–87. https://doi.org/10.1053/j.semdp.2023.02.002

Ali, T., & Muhammad, A. (2023). Artificial intelligence for antimicrobial resistance prediction: Challenges and opportunities towards practical implementation. Antibiotics, 12(3), Article 523. https://doi.org/10.3390/antibiotics12030523

Attia, Z. I., Noseworthy, P. A., Lopez-Jimenez, F., Asirvatham, S. J., Deshmukh, A. J., Gersh, B. J., Carter, R. E., & Friedman, P. A. (2019). An artificial intelligence-enabled ECG algorithm for the identification of patients with atrial fibrillation during sinus rhythm: A retrospective analysis of outcome prediction. The Lancet, 394(10201), 861–867. https://doi.org/10.1016/S0140-6736(19)31721-0

Augusto, A. C. C. N. (2025). Evaluating the robustness of AI-based vocal biomarkers against real-world noise: Toward regulatory standards and compliance [Master’s thesis, University of São Paulo].

Baddal, B., Taner, F., & Uzun Ozsahin, D. (2024). Harnessing of artificial intelligence for the diagnosis and prevention of hospital-acquired infections: A systematic review. Diagnostics, 14(5), Article 484. https://doi.org/10.3390/diagnostics14050484

Basem, J., Mani, R., Sun, S., Gilotra, K., Dianati-Maleki, N., & Dashti, R. (2025). Clinical applications of artificial intelligence and machine learning in neurocardiology: A comprehensive review. Frontiers in Cardiovascular Medicine, 12, Article 1525966. https://doi.org/10.3389/fcvm.2025.1525966

Bekbolatova, M., Mayer, J., Ong, C. W., & Toma, M. (2024). Transformative potential of AI in healthcare: Definitions, applications, and navigating the ethical landscape and public perspectives. Healthcare, 12(2), Article 125. https://doi.org/10.3390/healthcare12020125

Browning, L., Jesus, C., Malacrino, S., Guan, Y., White, K., Puddle, A., Alham, N. K., Haghighat, M., Colling, R., Birks, J., & Rittscher, J. (2024). Artificial intelligence-based quality assessment of histopathology whole-slide images within a clinical workflow: Assessment of ‘PathProfiler’ in a diagnostic pathology setting. Diagnostics, 14(10), Article 990. https://doi.org/10.3390/diagnostics14100990

Cau, R., Pisu, F., Porcu, M., Cademartiri, F., Montisci, R., Bassareo, P., & Saba, L. (2023). Machine learning approach in diagnosing Takotsubo cardiomyopathy: The role of the combined evaluation of atrial and ventricular strain, and parametric mapping. International Journal of Cardiology, 373, 124–133. https://doi.org/10.1016/j.ijcard.2022.11.021

Choi, J., Kim, J. Y., Cho, M. S., Kim, M., Kim, J., Oh, I. Y., & Kim, Y. H. (2024). Artificial intelligence predicts undiagnosed atrial fibrillation in patients with embolic stroke of undetermined source using sinus rhythm electrocardiograms. Heart Rhythm, 21(9), 1647–1655. https://doi.org/10.1016/j.hrthm.2024.03.029

Coudray, N., Ocampo, P. S., Sakellaropoulos, T., Narula, N., Snuderl, M., Fenyö, D., Moreira, A. L., Razavian, N., & Tsirigos, A. (2018). Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning. Nature Medicine, 24(10), 1559–1567. https://doi.org/10.1038/s41591-018-0177-5

Daly, J. E., Delen, D., Han, Z., Smith, R., Honerlaw, J., Cho, K., Bennett, B., & Sippel, J. (2026). AI in clinical decision support systems: Promising applications and strategies for managing data challenges. Journal of Medical Internet Research, 28, Article e71532. https://doi.org/10.2196/71532

Egli, A. (2023). ChatGPT, GPT-4, and other large language models: The next revolution for clinical microbiology? Clinical Infectious Diseases, 77(9), 1322–1328. https://doi.org/10.1093/cid/ciad407

Gador-Whyte, A. P., Sherry, N. L., Brischetto, A., Andersson, P., Bond, K. A., van Hal, S. J., Harris, P. N. A., & Howden, B. P. (2026). Implementation of pathogen genomics in clinical microbiology laboratories. Clinical Microbiology Reviews. Advance online publication.

Gasser, U. (2023). An EU landmark for AI governance. Science, 380(6651), 1203. https://doi.org/10.1126/science.adj2053

Giske, C. G., Bressan, M., Fiechter, F., Hinic, V., Mancini, S., Nolte, O., & Egli, A. (2024). GPT-4-based AI agents—The new expert system for detection of antimicrobial resistance mechanisms? Journal of Clinical Microbiology, 62(10), Article e0068924. https://doi.org/10.1128/jcm.00689-24

Guzman-Garcia, C., Mordonini, M., & Cagnoni, S. (2025). AI-driven applications for diagnostic, clinical decision support and prevention in cardiology, orthopedics, and oncology. In Proceedings of the Ital-IA 2025 Conference on Artificial Intelligence (CEUR Workshop Proceedings, Vol. 4121). CEUR-WS.org. https://ceur-ws.org/Vol-4121/Ital-IA_2025_paper_113.pdf

Hall, K. K., Sheedy, C., Shoemaker-Hunt, S., Wyant, B., Hoffman, L., Bacon, O., Richard, S., Hassol, A., Gall, E., Schneiderman, S., Schoyer, E., Woo, M., Costar, D., LeRoy, L., Gale, B., Fitall, E., Schiff, G., Long, A., Miller, K., & Lim, A. (2020). Making healthcare safer III: A critical analysis of existing and emerging patient safety practices (AHRQ Publication No. 20-0029-EF). Agency for Healthcare Research and Quality.

Hirosawa, T., & Shimizu, T. (2025). A narrative review of artificial intelligence in medical diagnostics. Computers, Materials & Continua, 83(3), 3919–3945. https://doi.org/10.32604/cmc.2025.063803

Laumer, F., Di Vece, D., Cammann, V. L., Würdinger, M., Petkova, V., Schönberger, M., & Templin, C. (2022). Assessment of artificial intelligence in echocardiography diagnostics in differentiating Takotsubo syndrome from myocardial infarction. JAMA Cardiology, 7(5), 494–503. https://doi.org/10.1001/jamacardio.2022.0183

Niu, Z., Kuang, X., Chen, J., Cai, X., & Zhang, P. (2025). The application and challenges of ChatGPT in laboratory medicine. Advances in Laboratory Medicine, 6(4), 385–396. https://doi.org/10.1515/almed-2025-0080

Sallam, M., Snygg, J., Allam, D., Kassem, R., & Damani, M. (2025). Artificial intelligence in clinical medicine: A SWOT analysis of AI progress in diagnostics, therapeutics, and safety. Journal of Innovations in Medical Research, 4(3), 1–20. https://doi.org/10.63593/jimr.2788-7022.2025.06.001

Schnetler, R., van der Vegt, A., Kalke, V. R., Lane, P., & Scott, I. (2025). False hope of a single generalisable AI sepsis prediction model: Bias and proposed mitigation strategies for improving performance based on a retrospective multisite cohort study. BMJ Quality & Safety, 34(9), 580–589. https://doi.org/10.1136/bmjqs-2024-018328

Serag, A., Ion-Margineanu, A., Qureshi, H., McMillan, R., Saint Martin, M. J., Diamond, J., O’Reilly, P., & Hamilton, P. (2019). Translational AI and deep learning in diagnostic pathology. Frontiers in Medicine, 6, Article 185. https://doi.org/10.3389/fmed.2019.00185

Shimizu, T., Hautz, W. E., van Sassen, C., & Zwaan, L. (2025). The global progress for improving diagnosis: What we’ve learned, what comes next. Diagnosis, 12(4), 529–537. https://doi.org/10.1515/dx-2025-0109

Singh, H., Meyer, A. N. D., & Thomas, E. J. (2014). The frequency of diagnostic errors in outpatient care: Estimations from three large observational studies involving US adult populations. BMJ Quality & Safety, 23(9), 727–731. https://doi.org/10.1136/bmjqs-2013-002627

Tricco, A. C., Lillie, E., Zarin, W., O’Brien, K. K., Colquhoun, H., Levac, D., Moher, D., Peters, M. D. J., Horsley, T., Weeks, L., Hempel, S., Akl, E. A., Chang, C., McGowan, J., Stewart, L., Hartling, L., Aldcroft, A., Wilson, M. G., Garritty, C., … Straus, S. E. (2018). PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Annals of Internal Medicine, 169(7), 467–473. https://doi.org/10.7326/M18-0850

Vandenberg, O., Durand, G., Hallin, M., Diefenbach, A., Gant, V., Murray, P., Kozlakidis, Z., & van Belkum, A. (2020). Consolidation of clinical microbiology laboratories and introduction of transformative technologies. Clinical Microbiology Reviews, 33(2), Article e00057-19. https://doi.org/10.1128/CMR.00057-19

Venturi, F., Veronesi, G., Gualandi, A., Magnaterra, E., Scotti, B., Sotiri, I., Baraldi, C., Alessandrini, A. M., Veneziano, L., Vaccari, S., & Dika, E. (2025). From slide to insight: The emerging alliance of digital pathology and AI in melanoma diagnostics. Cancers, 17(22), Article 3696. https://doi.org/10.3390/cancers17223696

Veronesi, G., Curti, N., Gardini, A., Querzoli, G., Castellani, G., & Dika, E. (2025). Machine learning to detect melanoma exploiting nuclei morphology and spatial organization. Scientific Reports, 15, Article 21594. https://doi.org/10.1038/s41598-025-21594-x

Wurcel, V., Cicchetti, A., Garrison, L., Kip, M. M. A., Koffijberg, H., Kolbe, A., Leeflang, M. M. G., Merlin, T., Mestre-Ferrandiz, J., Oortwijn, W., Oosterwijk, C., Tunis, S., & Zamora, B. (2019). The value of diagnostic information in personalised healthcare: A comprehensive concept to facilitate bringing this technology into healthcare systems. Public Health Genomics, 22(1–2), 8–15. https://doi.org/10.1159/000501832

Zhang, X., Liu, C., Sun, Y., You, L., Zhang, X., & Shang, H. (2026). Artificial intelligence medical diagnostic devices: A scoping review of clinical research, challenges and future directions. EngMedicine, 3(1), Article 100120. https://doi.org/10.1016/j.engmed.2026.100120