Integrative Biomedical Research
Integrative Biomedical Research (Journal of Angiotherapy) | Online ISSN 3068-6326
463
Citations
1.8m
Views
752
Articles
REVIEWS (Open Access)
Reproducibility, Algorithmic Bias, and Clinical Readiness in Medical Imaging AI: A Lifecycle-Based Synthesis and Multi-Institutional Validation Protocol
Mst Murshida Mahbub 1*, Rabiatul Basria S. M. N. Mydin 2, Chandrarohini Saravanan 2
Integrative Biomedical Research 10 (1) 1-8 https://doi.org/10.25163/biomedical.10110898
Submitted: 17 September 2026 Revised: 04 November 2026 Accepted: 11 November 2026 Published: 13 November 2026
Abstract
Deep learning has transformed diagnostic imaging, pathology, and ophthalmology, with the potential to reduce diagnostic errors and inter-observer variability that still burden routine radiology practice. However, translation of these models from laboratory benchmarks into everyday clinical workflows remains incomplete. This paper reports a structured narrative synthesis of the peer-reviewed and preprint literature addressing three interlocking evidence gaps in medical imaging artificial intelligence (AI): reproducibility failures, the multi-stage propagation of algorithmic bias, and the determinants of clinical readiness. Sources were organized around a five-stage AI lifecycle framework and cross-referenced against reporting frameworks including CLAIM, TRIPOD-AI, and the FUTURE-AI consensus guidelines. Across the cohorts reviewed, internally validated performance metrics (AUROC/DSC frequently exceeding 0.90) declined under external, multi-site testing in every case identified, with reported drops ranging from approximately 3 to 20 percentage points depending on task, modality, and cohort; these figures are drawn from heterogeneous studies using different metrics and are reported here as an illustrative pattern rather than a pooled effect size. Structural mitigation strategies-including Common Data Models, federated learning, and differential privacy-showed promise but remain unevenly adopted. Building on these gaps, we outline a prospective, three-arm multi-institutional protocol addressing reproducibility (RQ1), bias mitigation (RQ2), and clinical readiness (RQ3), intended as a template for future empirical validation rather than a completed study. Bridging the gap between technical proof-of-concept and safe clinical deployment will likely require simultaneous progress on standardized reporting, lifecycle-wide bias auditing, and governance structures capable of monitoring AI performance after initial approval.
Keywords: artificial intelligence, medical imaging, deep learning, reproducibility, algorithmic bias, external validation, clinical readiness, FUTURE-AI guidelines
References
Alabduljabbar, A., Khan, S. U., Alsuhaibani, A., Almarshad, F., & Altherwy, Y. N. (2024). Medical imaging datasets, preparation, and availability for artificial intelligence in medical imaging. Journal of Alzheimer's Disease Reports, 8(1), 1471-1483. https://doi.org/10.3233/ADR-240129
Alandejani, F., Alabdulkarim, B., Alaseri, M., Al-Ansari, M., Al-Dury, S., & Al-Mallah, M. H. (2022). Deep learning-based fully automatic segmentation of left and right ventricle from short-axis cine magnetic resonance images: Validation in multi-center clinical datasets. Journal of Cardiovascular Magnetic Resonance, 24(1), 25. https://doi.org/10.1186/s12968-022-00855-3
Allen, B., Jr., Seltzer, S. E., Langlotz, C. P., Dreyer, K. P., Summers, R. M., Petrick, N., Marinac-Dabic, D., Cruz, M., Alkasab, T. K., Hanisch, R. J., Nilsen, W. J., Burleson, J., Lyman, K., & Kandarpa, K. (2019). A road map for translational research on artificial intelligence in medical imaging: From the 2018 National Institutes of Health/RSNA/ACR/The Academy Workshop. Journal of the American College of Radiology, 16(9), 1179-1189. https://doi.org/10.1016/j.jacr.2019.04.011
Barberis, A., & Mezard, M. (2024). RENOIR: Repeated sampling methods for machine learning validation in biology and medicine. Scientific Reports, 14, 5116. https://doi.org/10.1038/s41598-024-51381-4
Bencevic, M., Habijan, M., Galic, I., Babin, D., & Pižurica, A. (2024). Understanding skin color bias in deep learning-based skin lesion segmentation. Computer Methods and Programs in Biomedicine, 245, Article 108044. https://doi.org/10.1016/j.cmpb.2024.108044
Cobo, M., Fontecha, D. C., Silva, W., & Iglesias, L. L. (2025). Preprocessing guidelines and FAIR4prep: Establishing best practices for clinical informatics workflows. Scientific Data, 12(1), 732. https://doi.org/10.1038/s41597-023-02641-x
Daye, D., Wiggins, W. F., Lungren, M. P., & Alkasab, T. K. (2022). Implementation of clinical artificial intelligence in radiology: A roadmap for governance, maintenance, and monitoring. Radiology, 305(3), 555-563. https://doi.org/10.1148/radiol.213123
Drukker, K., Chen, W., Gichoya, J., Gruszauskas, N., Kalpathy-Cramer, J., Koyejo, S., Myers, K., Sá, R. C., Sahiner, B., Whitney, H., Zhang, Z., & Giger, M. (2023). Toward fairness in artificial intelligence for medical image analysis: Identification and mitigation of potential biases in the roadmap from data collection to model deployment. Journal of Medical Imaging, 10(6), 061104. https://doi.org/10.1117/1.JMI.10.6.061104
Erickson, B. J., Khosravi, B., Vahdati, S., & Zhang, K. (2024). FDA review of radiologic AI algorithms: Process, cybersecurity, and clinical data curation challenges. Radiology, 310(2), e230242. https://doi.org/10.1148/ryai.230242
Fahad, N., Sadib, R. J., Sajib, R. H., Morol, M. K., Nandi, D., & Liew, T. H. (2026). Responsible artificial intelligence in medical imaging: A systematic review. Frontiers in Digital Health, 8, 1884692. https://doi.org/10.3389/fdgth.2026.1884692
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé, H., III, & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723
Gichoya, J. W., Banerjee, I., Bhimireddy, A. R., Burns, J. L., Celi, L. A., Chen, L. C., Correa, R., Dullerud, N., Ghassemi, M., Huang, S. C., Kuo, P. C., Lungren, M. P., Palmer, L. J., Price, B. J., Purkayastha, S., Pyrros, A. T., Oakden-Rayner, L., Okechukwu, C., Seyyed-Kalantari, L., & Zhang, H. (2022). AI recognition of patient race in medical imaging: A modelling study. The Lancet Digital Health, 4(6), e406-e414. https://doi.org/10.1016/S2589-7500(22)00063-2
Jeon, K., Park, W. Y., Schmidt, T. S., & You, S. C. (2026). Clinical data standardization for distributed research: Standardizing medical imaging data with Radiology and Medical Imaging Common Data Models (R-CDM and MI-CDM). Investigative Radiology, 61(Suppl), S74-S83. https://doi.org/10.1097/RLI.0000000000001155
Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in ML-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804
Keane, P. A., & Topol, E. J. (2018). With an eye to AI and autonomous diagnosis. npj Digital Medicine, 1, 40. [As cited in NCBI Bookshelf resource NBK619320]
Khor, S., Haupt, E. C., Hahn, E. E., Lyons, L. J., Shankaran, V., & Bansal, A. (2023). Racial and ethnic bias in risk prediction models for colorectal cancer recurrence when race and ethnicity are omitted as predictors. JAMA Network Open, 6(6), e2318495. https://doi.org/10.1001/jamanetworkopen.2023.18495
Kidwai-Khan, F., Wang, R., Skanderson, M., Brandt, C. A., Fodeh, S., & Womack, J. A. (2024). A roadmap to artificial intelligence (AI): Methods for designing and building AI ready data to promote fairness. Journal of Biomedical Informatics, 154, 104654. https://doi.org/10.1016/j.jbi.2024.104654
Kinahan, P. (2025). Centralized imaging collaborations and the landscape of medical imaging databases for artificial intelligence readiness. In Gilbert W. Beebe Symposium: AI and ML applications in radiation therapy, medical diagnostics, and radiation occupational health and safety (pp. 45-56). National Academies Press. https://doi.org/10.17226/29200
Kondylakis, H., Osuala, R., Puig-Bosch, X., Lazrak, N., Diaz, O., Kushibar, K., Chouvarda, I., Charalambous, S., Starmans, M. P., Colantonio, S., Tachos, N., Joshi, S., Woodruff, H. C., Salahuddin, Z., Tsakou, G., Aussó, S., Alberich, L. C., Papanikolaou, N., Lambin, P., Marias, K., Tsiknakis, M., Fotiadis, D. I., Martí-Bonmatí, L., & Lekadir, K. (2025). A review of methods for trustworthy AI in medical imaging: The FUTURE-AI guidelines. IEEE Journal of Biomedical and Health Informatics, 29(2), 2017-2026. https://doi.org/10.1109/JBHI.2025.3614546
Langlotz, C. P., Allen, B., Erickson, B. J., Kalpathy-Cramer, J., Bigelow, K., Cook, T. S., Flanders, A. E., Lungren, M. P., Mendelson, D. S., Rudie, J. D., & Wang, G. (2019). A roadmap for foundational research on artificial intelligence in medical imaging: From the 2018 NIH/RSNA/ACR/The Academy Workshop. Radiology, 291(3), 781-791. https://doi.org/10.1148/radiol.2019190613
Larrazabal, A. J., Nieto, N., Peterson, V., Milone, D. H., & Ferrante, E. (2020). Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proceedings of the National Academy of Sciences of the United States of America, 117(23), 12592-12594. https://doi.org/10.1073/pnas.1919012117
Lekadir, K., Frangi, A. F., Porras, A. R., Glocker, B., Cintas, C., Langlotz, C. P., Weicken, E., Asselbergs, F. W., Prior, F., Collins, G. S., Kaissis, G., Tsakou, G., Buvat, I., Kalpathy-Cramer, J., Mongan, J., Schnabel, J. A., Kushibar, K., Riklund, K., Marias, K., & FUTURE-AI Consortium. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ, 388, e081554. https://doi.org/10.1136/bmj-2024-081554
Lones, M. A. (2024). Avoiding common machine learning pitfalls. Patterns, 5(10), Article 101046. https://doi.org/10.1016/j.patter.2024.101046
Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., März, K., & Kopp-Schneider, A. (2018). Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. Nature Communications, 9, 5217. https://doi.org/10.1038/s41467-018-07619-7
Maier-Hein, L., Reinke, A., Kozubek, M., & BIAS Initiative. (2020). Good scientific practice in biomedical image analysis challenges: BIAS statement. Medical Image Analysis, 66, 101796. https://doi.org/10.1016/j.media.2020.101796
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, J., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220-229. https://doi.org/10.1145/3287560.3287596
Nooraie, R. Y., Shelton, R. C., Lee, M., Brotzman, L. E., & Gichoya, J. W. (2021). Applying an implementation science lens and equity focus to the scale-up of clinical artificial intelligence in medical imaging. PET Clinics, 16(4), 643-653. https://doi.org/10.1016/j.cpet.2021.07.002
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. https://doi.org/10.1126/science.aax2342
Ogut, E. (2025). Artificial intelligence in clinical medicine: Challenges across diagnostic imaging, clinical decision support, surgery, pathology, and drug discovery. Clinical Practice, 15(9), 169. https://doi.org/10.3390/clinpract15090169
Puyol-Antón, E., Ruijsink, B., Harana, J. M., Piechnik, S. K., Neubauer, S., Petersen, S. E., Razavi, R., & King, A. P. (2022). Fairness in cardiac magnetic resonance imaging: Assessing sex and racial bias in deep learning-based segmentation. Frontiers in Cardiovascular Medicine, 9, Article 859310. https://doi.org/10.3389/fcvm.2022.859310
Rädsch, T., Reinke, A., Weru, V., Kopp-Schneider, A., & Maier-Hein, L. (2023). Labeling instructions matter: Systematic evaluation of annotation noise and labeling standards in biomedical image analysis. Nature Machine Intelligence, 5, 254-267. https://doi.org/10.1038/s42256-023-00625-5
Rafique, S., Chaudhary, K., Haidar, S. H., Rashid, U., & Usman, S. (2026). Diagnostic AI across the life sciences (2015-2025): A PRISMA-scoping review and bibliometric synthesis of external validity, calibration, fairness, and reproducibility. Haya: The Saudi Journal of Life Sciences, 11(2), 122-141. https://doi.org/10.36348/sjls.2026.v11i02.002
Raposo, H. (2025). Artificial intelligence-enabled medical imaging for early disease detection: Methodological advances, clinical validation, and translational challenges. SSRN Electronic Journal, 5331997. https://doi.org/10.2139/ssrn.5331997
Rubak, M. W., Brejnebøl, M. H., Rose, M. H., Gudbergsen, H., Chaudhari, A., Troelsen, A., Moller, A., Nybing, J. U., & Boesen, M. (2025). Federated learning and robustness scenarios under label-noise, image-noise, and faulty-client conditions in medical imaging. Journal of Clinical Medicine, 14(3), 132-146.
Sen, C. K., & DeMazumder, D. (2025). Getting started on artificial intelligence in health care and clinical research: Includes rigor checklist for authors and reviewers. Journal of Wound Care, 34(2), 114-127. https://doi.org/10.12968/jowc.2025.34.2.114
Sulaimanov, U., Sanlier, N., Moniri, A., Demir, B., Serikkanov, Y., Bayramoglu, A. R., Al-Jebur, M. S., Uslu, I., Ozturk, O., Nizzola, M., Ötles, E., Ammanuel, S. G., Keles, A., Erginoglu, U., & Baskaya, M. K. (2026). Are AI neuroimaging models ready for clinical use? A systematic methodological review. Journal of Clinical Medicine, 15, 03441. https://doi.org/10.3390/jcm1503441
Szabo, L., Lekadir, K., Mosteiro, P., Mitchell, M., Goisauf, M., & Sardanelli, F. (2022). Developing a "trustworthy AI system" in medical imaging: Practical guidelines and questions. Frontiers in Cardiovascular Medicine, 9, Article 1016032. https://doi.org/10.3389/fcvm.2022.1016032
Tejani, A. S., Klontzas, M. E., Gatti, A. A., Mongan, J. T., Moy, L., Park, S. H., Kahn, C. E., & Panel, C. U. (2025). Methodological validation and reporting standards in AI-based diagnostic neuroradiological and neurosurgical research: A systematic review. Journal of Clinical Medicine, 15(9), 3441. https://doi.org/10.3390/jcm15093441
Young, A. T., Pfau, J., Keloth, V., Chande, D., Farzandipour, M., Al shareef, H. N., & Wei, M. L. (2020). Selective prediction and deep learning for robust medical image classification: A dermoscopy clinical reader study. npj Digital Medicine, 3, 4. https://doi.org/10.1038/s41746-020-00380-6
Yousefi Nooraie, R., Lyons, P. G., Baumann, A. A., & Saboury, B. (2025). Equitable implementation of artificial intelligence in medical imaging: What can be learned from implementation science? PET Clinics, 20(2), Article 682.
Recommended articles
Digital Transformation in Organizations: Determinants, Outcomes, and the Role of Artificial Intelligence—A Systematic Review
Article metrics
View details
0
Downloads
0
Citations
26
Views
0
Save
Save
0
Citation
Citation
26
View
View
0
Share
Share