Integrative Biomedical Research

Integrative Biomedical Research (Journal of Angiotherapy) | Online ISSN  3068-6326
463
Citations
1.8m
Views
752
Articles
Your new experience awaits. Try the new design now and help us make it even better
Switch to the new experience
REVIEWS   (Open Access)

Mst Murshida Mahbub 1*, Rabiatul Basria S. M. N. Mydin 2, Chandrarohini Saravanan 2

+ Author Affiliations

Integrative Biomedical Research 10 (1) 1-8 https://doi.org/10.25163/biomedical.10110898

Submitted: 17 September 2026 Revised: 04 November 2026  Accepted: 11 November 2026  Published: 13 November 2026 


Abstract

Deep learning has transformed diagnostic imaging, pathology, and ophthalmology, with the potential to reduce diagnostic errors and inter-observer variability that still burden routine radiology practice. However, translation of these models from laboratory benchmarks into everyday clinical workflows remains incomplete. This paper reports a structured narrative synthesis of the peer-reviewed and preprint literature addressing three interlocking evidence gaps in medical imaging artificial intelligence (AI): reproducibility failures, the multi-stage propagation of algorithmic bias, and the determinants of clinical readiness. Sources were organized around a five-stage AI lifecycle framework and cross-referenced against reporting frameworks including CLAIM, TRIPOD-AI, and the FUTURE-AI consensus guidelines. Across the cohorts reviewed, internally validated performance metrics (AUROC/DSC frequently exceeding 0.90) declined under external, multi-site testing in every case identified, with reported drops ranging from approximately 3 to 20 percentage points depending on task, modality, and cohort; these figures are drawn from heterogeneous studies using different metrics and are reported here as an illustrative pattern rather than a pooled effect size. Structural mitigation strategies-including Common Data Models, federated learning, and differential privacy-showed promise but remain unevenly adopted. Building on these gaps, we outline a prospective, three-arm multi-institutional protocol addressing reproducibility (RQ1), bias mitigation (RQ2), and clinical readiness (RQ3), intended as a template for future empirical validation rather than a completed study. Bridging the gap between technical proof-of-concept and safe clinical deployment will likely require simultaneous progress on standardized reporting, lifecycle-wide bias auditing, and governance structures capable of monitoring AI performance after initial approval.

Keywords: artificial intelligence, medical imaging, deep learning, reproducibility, algorithmic bias, external validation, clinical readiness, FUTURE-AI guidelines

References

Alabduljabbar, A., Khan, S. U., Alsuhaibani, A., Almarshad, F., & Altherwy, Y. N. (2024). Medical imaging datasets, preparation, and availability for artificial intelligence in medical imaging. Journal of Alzheimer's Disease Reports, 8(1), 1471-1483. https://doi.org/10.3233/ADR-240129

Alandejani, F., Alabdulkarim, B., Alaseri, M., Al-Ansari, M., Al-Dury, S., & Al-Mallah, M. H. (2022). Deep learning-based fully automatic segmentation of left and right ventricle from short-axis cine magnetic resonance images: Validation in multi-center clinical datasets. Journal of Cardiovascular Magnetic Resonance, 24(1), 25. https://doi.org/10.1186/s12968-022-00855-3

Allen, B., Jr., Seltzer, S. E., Langlotz, C. P., Dreyer, K. P., Summers, R. M., Petrick, N., Marinac-Dabic, D., Cruz, M., Alkasab, T. K., Hanisch, R. J., Nilsen, W. J., Burleson, J., Lyman, K., & Kandarpa, K. (2019). A road map for translational research on artificial intelligence in medical imaging: From the 2018 National Institutes of Health/RSNA/ACR/The Academy Workshop. Journal of the American College of Radiology, 16(9), 1179-1189. https://doi.org/10.1016/j.jacr.2019.04.011

Barberis, A., & Mezard, M. (2024). RENOIR: Repeated sampling methods for machine learning validation in biology and medicine. Scientific Reports, 14, 5116. https://doi.org/10.1038/s41598-024-51381-4

Bencevic, M., Habijan, M., Galic, I., Babin, D., & Pižurica, A. (2024). Understanding skin color bias in deep learning-based skin lesion segmentation. Computer Methods and Programs in Biomedicine, 245, Article 108044. https://doi.org/10.1016/j.cmpb.2024.108044

Cobo, M., Fontecha, D. C., Silva, W., & Iglesias, L. L. (2025). Preprocessing guidelines and FAIR4prep: Establishing best practices for clinical informatics workflows. Scientific Data, 12(1), 732. https://doi.org/10.1038/s41597-023-02641-x

Daye, D., Wiggins, W. F., Lungren, M. P., & Alkasab, T. K. (2022). Implementation of clinical artificial intelligence in radiology: A roadmap for governance, maintenance, and monitoring. Radiology, 305(3), 555-563. https://doi.org/10.1148/radiol.213123

Drukker, K., Chen, W., Gichoya, J., Gruszauskas, N., Kalpathy-Cramer, J., Koyejo, S., Myers, K., Sá, R. C., Sahiner, B., Whitney, H., Zhang, Z., & Giger, M. (2023). Toward fairness in artificial intelligence for medical image analysis: Identification and mitigation of potential biases in the roadmap from data collection to model deployment. Journal of Medical Imaging, 10(6), 061104. https://doi.org/10.1117/1.JMI.10.6.061104

Erickson, B. J., Khosravi, B., Vahdati, S., & Zhang, K. (2024). FDA review of radiologic AI algorithms: Process, cybersecurity, and clinical data curation challenges. Radiology, 310(2), e230242. https://doi.org/10.1148/ryai.230242

Fahad, N., Sadib, R. J., Sajib, R. H., Morol, M. K., Nandi, D., & Liew, T. H. (2026). Responsible artificial intelligence in medical imaging: A systematic review. Frontiers in Digital Health, 8, 1884692. https://doi.org/10.3389/fdgth.2026.1884692

Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé, H., III, & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723

Gichoya, J. W., Banerjee, I., Bhimireddy, A. R., Burns, J. L., Celi, L. A., Chen, L. C., Correa, R., Dullerud, N., Ghassemi, M., Huang, S. C., Kuo, P. C., Lungren, M. P., Palmer, L. J., Price, B. J., Purkayastha, S., Pyrros, A. T., Oakden-Rayner, L., Okechukwu, C., Seyyed-Kalantari, L., & Zhang, H. (2022). AI recognition of patient race in medical imaging: A modelling study. The Lancet Digital Health, 4(6), e406-e414. https://doi.org/10.1016/S2589-7500(22)00063-2

Jeon, K., Park, W. Y., Schmidt, T. S., & You, S. C. (2026). Clinical data standardization for distributed research: Standardizing medical imaging data with Radiology and Medical Imaging Common Data Models (R-CDM and MI-CDM). Investigative Radiology, 61(Suppl), S74-S83. https://doi.org/10.1097/RLI.0000000000001155

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in ML-based science. Patterns, 4(9), Article 100804. https://doi.org/10.1016/j.patter.2023.100804

Keane, P. A., & Topol, E. J. (2018). With an eye to AI and autonomous diagnosis. npj Digital Medicine, 1, 40. [As cited in NCBI Bookshelf resource NBK619320]

Khor, S., Haupt, E. C., Hahn, E. E., Lyons, L. J., Shankaran, V., & Bansal, A. (2023). Racial and ethnic bias in risk prediction models for colorectal cancer recurrence when race and ethnicity are omitted as predictors. JAMA Network Open, 6(6), e2318495. https://doi.org/10.1001/jamanetworkopen.2023.18495

Kidwai-Khan, F., Wang, R., Skanderson, M., Brandt, C. A., Fodeh, S., & Womack, J. A. (2024). A roadmap to artificial intelligence (AI): Methods for designing and building AI ready data to promote fairness. Journal of Biomedical Informatics, 154, 104654. https://doi.org/10.1016/j.jbi.2024.104654

Kinahan, P. (2025). Centralized imaging collaborations and the landscape of medical imaging databases for artificial intelligence readiness. In Gilbert W. Beebe Symposium: AI and ML applications in radiation therapy, medical diagnostics, and radiation occupational health and safety (pp. 45-56). National Academies Press. https://doi.org/10.17226/29200

Kondylakis, H., Osuala, R., Puig-Bosch, X., Lazrak, N., Diaz, O., Kushibar, K., Chouvarda, I., Charalambous, S., Starmans, M. P., Colantonio, S., Tachos, N., Joshi, S., Woodruff, H. C., Salahuddin, Z., Tsakou, G., Aussó, S., Alberich, L. C., Papanikolaou, N., Lambin, P., Marias, K., Tsiknakis, M., Fotiadis, D. I., Martí-Bonmatí, L., & Lekadir, K. (2025). A review of methods for trustworthy AI in medical imaging: The FUTURE-AI guidelines. IEEE Journal of Biomedical and Health Informatics, 29(2), 2017-2026. https://doi.org/10.1109/JBHI.2025.3614546

Langlotz, C. P., Allen, B., Erickson, B. J., Kalpathy-Cramer, J., Bigelow, K., Cook, T. S., Flanders, A. E., Lungren, M. P., Mendelson, D. S., Rudie, J. D., & Wang, G. (2019). A roadmap for foundational research on artificial intelligence in medical imaging: From the 2018 NIH/RSNA/ACR/The Academy Workshop. Radiology, 291(3), 781-791. https://doi.org/10.1148/radiol.2019190613

Larrazabal, A. J., Nieto, N., Peterson, V., Milone, D. H., & Ferrante, E. (2020). Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proceedings of the National Academy of Sciences of the United States of America, 117(23), 12592-12594. https://doi.org/10.1073/pnas.1919012117

Lekadir, K., Frangi, A. F., Porras, A. R., Glocker, B., Cintas, C., Langlotz, C. P., Weicken, E., Asselbergs, F. W., Prior, F., Collins, G. S., Kaissis, G., Tsakou, G., Buvat, I., Kalpathy-Cramer, J., Mongan, J., Schnabel, J. A., Kushibar, K., Riklund, K., Marias, K., & FUTURE-AI Consortium. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ, 388, e081554. https://doi.org/10.1136/bmj-2024-081554

Lones, M. A. (2024). Avoiding common machine learning pitfalls. Patterns, 5(10), Article 101046. https://doi.org/10.1016/j.patter.2024.101046

Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., März, K., & Kopp-Schneider, A. (2018). Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. Nature Communications, 9, 5217. https://doi.org/10.1038/s41467-018-07619-7

Maier-Hein, L., Reinke, A., Kozubek, M., & BIAS Initiative. (2020). Good scientific practice in biomedical image analysis challenges: BIAS statement. Medical Image Analysis, 66, 101796. https://doi.org/10.1016/j.media.2020.101796

Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, J., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency, 220-229. https://doi.org/10.1145/3287560.3287596

Nooraie, R. Y., Shelton, R. C., Lee, M., Brotzman, L. E., & Gichoya, J. W. (2021). Applying an implementation science lens and equity focus to the scale-up of clinical artificial intelligence in medical imaging. PET Clinics, 16(4), 643-653. https://doi.org/10.1016/j.cpet.2021.07.002

Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. https://doi.org/10.1126/science.aax2342

Ogut, E. (2025). Artificial intelligence in clinical medicine: Challenges across diagnostic imaging, clinical decision support, surgery, pathology, and drug discovery. Clinical Practice, 15(9), 169. https://doi.org/10.3390/clinpract15090169

Puyol-Antón, E., Ruijsink, B., Harana, J. M., Piechnik, S. K., Neubauer, S., Petersen, S. E., Razavi, R., & King, A. P. (2022). Fairness in cardiac magnetic resonance imaging: Assessing sex and racial bias in deep learning-based segmentation. Frontiers in Cardiovascular Medicine, 9, Article 859310. https://doi.org/10.3389/fcvm.2022.859310

Rädsch, T., Reinke, A., Weru, V., Kopp-Schneider, A., & Maier-Hein, L. (2023). Labeling instructions matter: Systematic evaluation of annotation noise and labeling standards in biomedical image analysis. Nature Machine Intelligence, 5, 254-267. https://doi.org/10.1038/s42256-023-00625-5

Rafique, S., Chaudhary, K., Haidar, S. H., Rashid, U., & Usman, S. (2026). Diagnostic AI across the life sciences (2015-2025): A PRISMA-scoping review and bibliometric synthesis of external validity, calibration, fairness, and reproducibility. Haya: The Saudi Journal of Life Sciences, 11(2), 122-141. https://doi.org/10.36348/sjls.2026.v11i02.002

Raposo, H. (2025). Artificial intelligence-enabled medical imaging for early disease detection: Methodological advances, clinical validation, and translational challenges. SSRN Electronic Journal, 5331997. https://doi.org/10.2139/ssrn.5331997

Rubak, M. W., Brejnebøl, M. H., Rose, M. H., Gudbergsen, H., Chaudhari, A., Troelsen, A., Moller, A., Nybing, J. U., & Boesen, M. (2025). Federated learning and robustness scenarios under label-noise, image-noise, and faulty-client conditions in medical imaging. Journal of Clinical Medicine, 14(3), 132-146.

Sen, C. K., & DeMazumder, D. (2025). Getting started on artificial intelligence in health care and clinical research: Includes rigor checklist for authors and reviewers. Journal of Wound Care, 34(2), 114-127. https://doi.org/10.12968/jowc.2025.34.2.114

Sulaimanov, U., Sanlier, N., Moniri, A., Demir, B., Serikkanov, Y., Bayramoglu, A. R., Al-Jebur, M. S., Uslu, I., Ozturk, O., Nizzola, M., Ötles, E., Ammanuel, S. G., Keles, A., Erginoglu, U., & Baskaya, M. K. (2026). Are AI neuroimaging models ready for clinical use? A systematic methodological review. Journal of Clinical Medicine, 15, 03441. https://doi.org/10.3390/jcm1503441

Szabo, L., Lekadir, K., Mosteiro, P., Mitchell, M., Goisauf, M., & Sardanelli, F. (2022). Developing a "trustworthy AI system" in medical imaging: Practical guidelines and questions. Frontiers in Cardiovascular Medicine, 9, Article 1016032. https://doi.org/10.3389/fcvm.2022.1016032

Tejani, A. S., Klontzas, M. E., Gatti, A. A., Mongan, J. T., Moy, L., Park, S. H., Kahn, C. E., & Panel, C. U. (2025). Methodological validation and reporting standards in AI-based diagnostic neuroradiological and neurosurgical research: A systematic review. Journal of Clinical Medicine, 15(9), 3441. https://doi.org/10.3390/jcm15093441

Young, A. T., Pfau, J., Keloth, V., Chande, D., Farzandipour, M., Al shareef, H. N., & Wei, M. L. (2020). Selective prediction and deep learning for robust medical image classification: A dermoscopy clinical reader study. npj Digital Medicine, 3, 4. https://doi.org/10.1038/s41746-020-00380-6

Yousefi Nooraie, R., Lyons, P. G., Baumann, A. A., & Saboury, B. (2025). Equitable implementation of artificial intelligence in medical imaging: What can be learned from implementation science? PET Clinics, 20(2), Article 682.


Article metrics
View details
0
Downloads
0
Citations
26
Views

View Dimensions


View Plumx


View Altmetric



0
Save
0
Citation
26
View
0
Share