Article Contents
REVIEW   Open Access     Cite

Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI

    Show all affliationsShow less
More Information
  • DownLoad: Full size image
    1. Sustainable medical artificial intelligence development relies on robust statistical design and inference.

      Statistics address key medical AI issues: limited interpretability, overfitting, AI "hallucinations", etc.

      Randomized controlled trials are key to validating medical AI systems before clinical implementation.

  • In the rapidly evolving field of artificial intelligence (AI), statistics plays a crucial role in addressing challenges faced by medical AI. This review begins by highlighting the primary tasks of medical AI and the integration of statistical methodologies into their modeling processes. Despite the widespread application of AI in medicine and healthcare, key challenges persist: poor model interpretability, lack of causal reasoning, overfitting, unfairness, imbalanced dataset, AI "hallucinations" and "disinformation". Statistics provides unique strategies to tackle these challenges, including rigorous statistical design, regularization techniques, and statistical frameworks grounded in causal inference. Finally, the review offers several recommendations for the sustainable development of medical AI: enhancing data quality, promoting model simplicity and transparency, fostering independent validation standards, and facilitating interdisciplinary collaboration between statisticians and medical AI practitioners.
  • 加载中
  • [1] Nilsson N. J. (2009). The Quest for Artificial Intelligence (Cambridge University Press). DOI: 10.1017/CBO9780511819346. https://www.cambridge.org/core/product/32C727961B24223BBB1B3511F44F343E.

    View in Article Google Scholar

    [2] Feigenbaum E. A. (1977). The art of artificial intelligence: Themes and case studies of knowledge engineering. International Joint Conference on Artificial Intelligence. DOI:10.5555/892148.

    View in Article Google Scholar

    [3] Michalski R. S., Carbonell J. G. and Mitchell T. M. (1985). Machine learning: An artificial intelligence approach. Artif. Intell. 25:236−238. DOI:10.1016/0004-3702(85)90005-0

    View in Article CrossRef Google Scholar

    [4] Biffi C., Cerrolaza J. J., Tarroni G., et al. (2020). Explainable anatomical shape analysis through deep hierarchical generative models. IEEE Trans. Med. Imaging 39:2088−2099. DOI:10.1109/TMI.2020.2964499

    View in Article CrossRef Google Scholar

    [5] Brown T. B., Mann B., Ryder N., et al. (2020). Language models are few-shot learners. Proceedings of the 34th International Conference on Neural Information Processing Systems:Article 159. DOI:10.48550/arXiv.2005.14165

    View in Article Google Scholar

    [6] Topol E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nat. Med. 25:44−56. DOI:10.1038/s41591-018-0300-7

    View in Article CrossRef Google Scholar

    [7] Esteva A., Kuprel B., Novoa R. A., et al. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature 542:115−118. DOI:10.1038/nature21056

    View in Article CrossRef Google Scholar

    [8] Huff D. T., Weisman A. J. and Jeraj R. (2021). Interpretation and visualization techniques for deep learning models in medical imaging. Phys. Med. Biol. 66:04tr01. DOI:10.1088/1361-6560/abcd17

    View in Article CrossRef Google Scholar

    [9] Miotto R., Li L., Kidd B. A., et al. (2016). Deep patient: An unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6:26094. DOI:10.1038/srep26094

    View in Article CrossRef Google Scholar

    [10] Li Y., Xu Y., Möller J., et al. (2020). Life-course blood pressure trajectories and cardiovascular diseases: A population-based cohort study in China. Plos One 15:e0240804. DOI:10.1371/journal.pone.0240804

    View in Article CrossRef Google Scholar

    [11] Liu Y., Logan B., Liu N., et al. (2017). Deep reinforcement learning for dynamic treatment regimes on medical registry data. 2017 IEEE International Conference on Healthcare Informatics (ICHI) pp:380-385. DOI:10.1109/ICHI.2017.45

    View in Article Google Scholar

    [12] Bzdok D. (2017). Classical statistics and statistical learning in imaging neuroscience. Front. Neurosci. 11:543. DOI:10.3389/fnins.2017.00543

    View in Article CrossRef Google Scholar

    [13] Jordan M. I. and Mitchell T. M. (2015). Machine learning: Trends, perspectives, and prospects. Science 349:255−260. DOI:10.1126/science.aaa8415

    View in Article CrossRef Google Scholar

    [14] Hastie T., Tibshirani R., Friedman J., et al. (2004). The elements of statistical learning: Data mining, inference, and prediction. Math. Intell. 27:83−85. DOI:10.1007/BF02985802

    View in Article CrossRef Google Scholar

    [15] Murphy K. P. (2012). Machine learning: a probabilistic perspective (MIT press). https://www.cs.ubc.ca/~murphyk/MLbook/pml-toc-1may12.pdf.

    View in Article Google Scholar

    [16] Sokolova M. and Lapalme G. (2009). A systematic analysis of performance measures for classification tasks. Inform. Process. Manag. 45:427−437. DOI:10.1016/j.ipm.2009.03.002

    View in Article CrossRef Google Scholar

    [17] Bishop C. M. and Nasrabadi N. M. (2006). Pattern recognition and machine learning (Springer). https://link.springer.com/book/9780387310732.

    View in Article Google Scholar

    [18] LeCun Y., Bengio Y. and Hinton G. (2015). Deep learning. Nature 521:436−444. DOI:10.1038/nature14539

    View in Article CrossRef Google Scholar

    [19] Gelman A., Carlin J. B., Stern H. S., et al. (1995). Bayesian data analysis (Chapman and Hall/CRC). https://sites.stat.columbia.edu/gelman/book/BDA3.pdf.

    View in Article Google Scholar

    [20] Goodfellow I., Bengio Y. and Courville A. (2016). Deep learning (The MIT press). https://mitpress.mit.edu/9780262035613/deep-learning/.

    View in Article Google Scholar

    [21] Pan S. J. and Yang Q. (2009). A survey on transfer learning. IEEE Trans. Knowl. Data Eng. 22:1345−1359. DOI:10.1109/tkde.2009.191

    View in Article CrossRef Google Scholar

    [22] Hastie T., Tibshirani R., Friedman J. (2009). The elements of statistical learning: Data mining, inference, and prediction (Springer). https://link.springer.com/book/10.1007/978-0-387-84858-7.

    View in Article Google Scholar

    [23] Qi L., Zhang J., Qi Z.-F., et al. (2023). Measurement and evaluation method of radar anti-jamming effectiveness based on principal component analysis and machine learning. EURASIP Journal on Wireless Communications and Networking 2023:58. DOI:10.1186/s13638-023-02262-3

    View in Article CrossRef Google Scholar

    [24] John Lu Z. Q. (2010). Bayesian methods for data analysis, third edition. J. Appl. Stat. 37:705−706. DOI:10.1080/02664760902811621

    View in Article CrossRef Google Scholar

    [25] Maharana K., Mondal S. and Nemade B. (2022). A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings 3:91−99. DOI:10.1016/j.gltp.2022.04.020

    View in Article CrossRef Google Scholar

    [26] Wang Q., Reps J. M., Kostka K. F., et al. (2020). Development and validation of a prognostic model predicting symptomatic hemorrhagic transformation in acute ischemic stroke at scale in the OHDSI network. PLoS One 15:e0226718. DOI:10.1371/journal.pone.0226718

    View in Article CrossRef Google Scholar

    [27] Liu H. and Dong G. (2018). Feature engineering for machine learning and data analytics. (CRC Press). DOI:10.1201/9781315181080

    View in Article Google Scholar

    [28] Demirtas H. (2018). Flexible imputation of missing data. J. Stat. Softw. 85:1−5. DOI:10.18637/jss.v085.b04

    View in Article CrossRef Google Scholar

    [29] Kuhn M. and Johnson K. (2013). Applied Predictive Modeling. (Springer Publishing Company, Incorporated). DOI:10.1007/978-1-4614-6849-3

    View in Article Google Scholar

    [30] Pedregosa F., Varoquaux G., Gramfort A., et al. (2011). Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 12:2825−2830. DOI:10.48550/arXiv.1201.0490

    View in Article CrossRef Google Scholar

    [31] Jensen F. and Nielsen T. (2001). Bayesian networks and decision graphs. (Springer Publishing Company, Incorporated). DOI:10.1007/978-1-4757-3502-4

    View in Article Google Scholar

    [32] Jurafsky D. and Martin J. (2008). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. (Prentice Hall Press). https://pages.ucsd.edu/~bakovic/compphon/Jurafsky,%20Martin.-Speech%20and%20Language%20Processing_%20An%20Introduction%20to%20Natural%20Language%20Processing%20(2007).pdf.

    View in Article Google Scholar

    [33] Snoek J., Larochelle H. and Adams R. (2012). Practical bayesian optimization of machine learning algorithms. Advances in Neural Information Processing Systems 4. DOI:10.48550/arXiv.1206.2944

    View in Article Google Scholar

    [34] Stone M. (1974). Cross-validatory choice and assessment of statistical predictions. Journal of the Royal Statistical Society: Series B (Methodological) 36:111−133. DOI:10.1111/j.2517-6161.1974.tb00994.x

    View in Article CrossRef Google Scholar

    [35] Powers D. and Ailab (2011). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness & correlation. J. Mach. Learn. Technol. 2:2229-3981. DOI:10.9735/2229-3981

    View in Article Google Scholar

    [36] Hutter F., Kotthoff L. and Vanschoren J. (2019). Automated machine learning: Methods, systems, challenges (Springer Publishing Company, Incorporated). DOI:10.1007/978-3-030-05318-5

    View in Article Google Scholar

    [37] Sengar S. S., Hasan A. B., Kumar S., et al. (2024). Generative artificial intelligence: A systematic review and applications. Machine Learning arXiv:2405.11029. DOI:10.48550/arXiv.2405.11029

    View in Article Google Scholar

    [38] Dhoni P. S. (2023). Exploring the synergy between generative AI, data and analytics in the modern age. Authorea Preprints. https://www.techrxiv.org/users/691566/articles/682737-exploring-the-synergy-between-generative-ai-data-and-analytics-in-the-modern-age.

    View in Article Google Scholar

    [39] Adams S. and Beling P. A. (2019). A survey of feature selection methods for Gaussian mixture models and hidden Markov models. Artif. Intell. Rev. 52:1739−1779. DOI:10.1007/s10462-017-9581-3

    View in Article CrossRef Google Scholar

    [40] Kalingeri V. (2022). Latent variable modelling using variational autoencoders: A survey. Machine Learning arXiv:2206.09891. DOI:10.48550/arXiv.2206.09891

    View in Article Google Scholar

    [41] Rizvi S. K. J., Azad M. A. and Fraz M. M. (2021). Spectrum of advancements and developments in multidisciplinary domains for generative adversarial networks (GANs). Arch. Comput. Methods Eng. 28:4503−4521. DOI:10.1007/s11831-021-09543-4

    View in Article CrossRef Google Scholar

    [42] Jiang K. and Huang J. (2024). A Survey on vision autoregressive model. arXiv preprint arXiv:2411.08666. DOI:10.48550/2411.08666

    View in Article Google Scholar

    [43] Kondurkar I., Raj A. and Lakshmi D. (2024). Modern applications with a focus on training ChatGPT and GPT models: Exploring generative AI and NLP. Obaid, A. J., et al (Eds). Advanced Applications of Generative AI and Natural Language Processing Models (IGI Global) pp:186-227. DOI:10.4018/979-8-3693-0502-7.ch010

    View in Article Google Scholar

    [44] Chapman A. D. (2023). Artificial Intelligence and Machine Learning (The Autodidact’s Toolkit).

    View in Article Google Scholar

    [45] Bacanin N., Stoean R., Zivkovic M., et al. (2021). Performance of a novel chaotic firefly algorithm with enhanced exploration for tackling global optimization problems: Application for dropout regularization. Mathematics 9:2705. DOI:10.3390/math9212705

    View in Article CrossRef Google Scholar

    [46] Wang H. and Yeung D.-Y. (2020). A survey on Bayesian deep learning. arXiv preprint 53:1−37. DOI:10.48550/arXiv.1604.01662

    View in Article CrossRef Google Scholar

    [47] Dai J., Xu H., Chen T., et al. (2025). Artificial intelligence for medicine 2025: Navigating the endless frontier. Innov. Med. 3:100120. DOI:10.59717/j.xinn-med.2025.100120

    View in Article CrossRef Google Scholar

    [48] Tang Y.-D., Dong E.-D. and Gao W. (2024). LLMs in medicine: The need for advanced evaluation systems for disruptive technologies. The Innovation 5:100622. DOI:10.1016/j.xinn.2024.100622

    View in Article CrossRef Google Scholar

    [49] Huang T., Xu H., Wang H., et al. (2023). Artificial intelligence for medicine: Progress, challenges, and perspectives. Innov. Med. 1:100030. DOI:10.59717/j.xinn-med.2023.100030

    View in Article CrossRef Google Scholar

    [50] Announcement H., Memo H., Statement C. R. O., et al. (2023). Health Subcommittee Hearing: "Understanding How AI is Changing Health Care". https://energycommerce.house.gov/events/health-subcommittee-hearing-understanding-how-ai-is-changing-health-care.

    View in Article Google Scholar

    [51] House T. W. (2023). Delivering on the Promise of AI to Improve Health Outcomes. https://www.whitehouse.gov/briefing-room/blog/2023/12/14/delivering-on-the-promise-of-ai-to-improve-health-outcomes/.

    View in Article Google Scholar

    [52] Commission N. H. (2024). http://www.nhc.gov.cn/guihuaxxs/gongwen12/202411/647062ee76764323b29a1f0124b64400.shtml.

    View in Article Google Scholar

    [53] Pan Z., Zhang R., Shen S., et al. (2023). OWL: An optimized and independently validated machine learning prediction model for lung cancer screening based on the UK Biobank, PLCO, and NLST populations. EBioMedicine 88:104443. DOI:10.1016/j.ebiom.2023.104443

    View in Article CrossRef Google Scholar

    [54] Thirunavukarasu A. J., Ting D. S. J., Elangovan K., et al. (2023). Large language models in medicine. Nat. Med. 29:1930−1940. DOI:10.1038/s41591-023-02448-8

    View in Article CrossRef Google Scholar

    [55] Kraljevic Z., Bean D., Shek A., et al. (2022). Foresight--generative pretrained transformer (GPT) for modelling of patient timelines using Ehrs. arXiv preprint arXiv:2212.08072. DOI:10.48550/arXiv.2212.08072

    View in Article Google Scholar

    [56] Mayourian J., El-Bokl A., Lukyanenko P., et al. (2024). Electrocardiogram-based deep learning to predict mortality in paediatric and adult congenital heart disease. Eur. Heart J. 46:856−868. DOI:10.1093/eurheartj/ehae651

    View in Article CrossRef Google Scholar

    [57] Vamathevan J., Clark D., Czodrowski P., et al. (2019). Applications of machine learning in drug discovery and development. Nat. Rev. Drug Discovery 18:463−477. DOI:10.1038/s41573-019-0024-5

    View in Article CrossRef Google Scholar

    [58] Rohl C. A., Strauss C. E., Misura K. M., et al. (2004). Protein structure prediction using Rosetta. Methods Enzymol. 383:66−93. DOI:10.1016/s0076-6879(04)83004-0

    View in Article CrossRef Google Scholar

    [59] Abramson J., Adler J., Dunger J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630:493−500. DOI:10.1038/s41586-024-07487-w

    View in Article CrossRef Google Scholar

    [60] AlQuraishi M. (2019). AlphaFold at CASP13. Bioinformatics 35:4862−4865. DOI:10.1093/bioinformatics/btz422

    View in Article CrossRef Google Scholar

    [61] Dauparas J., Anishchenko I., Bennett N., et al. (2022). Robust deep learning–based protein sequence design using ProteinMPNN. Science 378:49−56. DOI:10.1126/science.add2187

    View in Article CrossRef Google Scholar

    [62] Zhavoronkov A., Ivanenkov Y. A., Aliper A., et al. (2019). Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat. Biotechnol. 37:1038−1040. DOI:10.1038/s41587-019-0224-x

    View in Article CrossRef Google Scholar

    [63] Tu T., Schaekermann M., Palepu A., et al. (2025). Towards conversational diagnostic artificial intelligence. Nature 642:442−450. DOI:10.1038/s41586-025-08866-7

    View in Article CrossRef Google Scholar

    [64] Waisberg E., Ong J., Masalkhi M., et al. (2023). GPT-4: A new era of artificial intelligence in medicine. Ir. J. Med. Sci. 192:3197−3200. DOI:10.1007/s11845-023-03377-8

    View in Article CrossRef Google Scholar

    [65] Singhal K., Tu T., Gottweis J., et al. (2025). Toward expert-level medical question answering with large language models. Nat. Med. 31:943−950. DOI:10.1038/s41591-024-03423-7

    View in Article CrossRef Google Scholar

    [66] Li Y., Li Z., Zhang K., et al. (2023). Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus 15:e40895. DOI:10.7759/cureus.40895

    View in Article CrossRef Google Scholar

    [67] Han T., Adams L., Papaioannou M., et al. (2023). MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data. arXiv preprint arXiv:2304.08247. DOI:10.48550/arXiv.2304.08247

    View in Article Google Scholar

    [68] Wu C., Zhang X., Zhang Y., et al. (2023). Pmc-llama: Further finetuning llama on medical papers. arXiv preprint 2:6. DOI:10.48550/arXiv.2304.14454

    View in Article CrossRef Google Scholar

    [69] Wang H., Liu C., Xi N., et al. (2023). Huatuo: Tuning llama model with chinese medical knowledge. arXiv preprint arXiv:2304.06975. DOI:10.48550/arXiv.2304.06975

    View in Article Google Scholar

    [70] Toma A., Lawler P. R., Ba J., et al. (2023). Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding. arXiv preprint arXiv:2305.1203. DOI:10.48550/arXiv.2305.1203

    View in Article Google Scholar

    [71] Liu X., Liu H., Yang G., et al. (2025). A generalist medical language model for disease diagnosis assistance. Nat. Med. 31:932−942. DOI:10.1038/s41591-024-03416-6

    View in Article CrossRef Google Scholar

    [72] Patel S. B. and Lam K. (2023). ChatGPT: The future of discharge summaries. Lancet Digit. Health 5:e107−e108. DOI:10.1016/S2589-7500(23)00021-3

    View in Article CrossRef Google Scholar

    [73] Hall A., Mitchell A. R. J., Wood L., et al. (2020). Effectiveness of a single lead AliveCor electrocardiogram application for the screening of atrial fibrillation: A systematic review. Medicine (Baltimore) 99:e21388. DOI:10.1097/MD.0000000000021388

    View in Article CrossRef Google Scholar

    [74] Abraham S. B., Arunachalam S., Zhong A., et al. (2021). Improved real-world glycemic control with continuous glucose monitoring system predictive alerts. J. Diabetes Sci. Technol. 15:91−97. DOI:10.1177/1932296819859334

    View in Article CrossRef Google Scholar

    [75] Li J., Guan Z., Wang J., et al. (2024). Integrated image-based deep learning and language models for primary diabetes care. Nat. Med. 30:2886−2896. DOI:10.1038/s41591-024-03139-8

    View in Article CrossRef Google Scholar

    [76] Subramanian M., Wojtusciszyn A., Favre L., et al. (2020). Precision medicine in the era of artificial intelligence: implications in chronic disease management. J. Transl. Med. 18:472. DOI:10.1186/s12967-020-02658-5

    View in Article CrossRef Google Scholar

    [77] Mendes-Soares H., Raveh-Sadka T., Azulay S., et al. (2019). Assessment of a personalized approach to predicting postprandial glycemic responses to food among individuals without diabetes. JAMA Netw. Open 2:e188102. DOI:10.1001/jamanetworkopen.2018.8102

    View in Article CrossRef Google Scholar

    [78] Xie Y., Seth I., Hunter-Smith D. J., et al. (2023). Aesthetic surgery advice and counseling from artificial intelligence: A rhinoplasty consultation with ChatGPT. Aesthetic Plast. Surg. 47:1985−1993. DOI:10.1007/s00266-023-03338-7

    View in Article CrossRef Google Scholar

    [79] Montazeri M., Multmeier J., Novorol C., et al. (2021). Optimization of patient flow in urgent care centers using a digital tool for recording patient symptoms and history: Simulation study. JMIR Form. Res. 5:e26402. DOI:10.2196/26402

    View in Article CrossRef Google Scholar

    [80] Cotte F., Mueller T., Gilbert S., et al. (2022). Safety of triage self-assessment using a symptom assessment app for walk-in patients in the emergency care setting: Observational prospective cross-sectional study. JMIR Mhealth Uhealth 10:e32340. DOI:10.2196/32340

    View in Article CrossRef Google Scholar

    [81] Fraser H. S. F., Cohan G., Koehler C., et al. (2022). Evaluation of diagnostic and triage accuracy and usability of a symptom checker in an emergency department: Observational study. JMIR Mhealth Uhealth 10:e38364. DOI:10.2196/38364

    View in Article CrossRef Google Scholar

    [82] Gilbert S., Mehl A., Baluch A., et al. (2020). How accurate are digital symptom assessment apps for suggesting conditions and urgency advice. A clinical vignettes comparison to GPs. BMJ Open 10:e040269. DOI:10.1136/bmjopen-2020-040269

    View in Article CrossRef Google Scholar

    [83] health A. (2025). Health. Powered by Ada. https://ada.com/.

    View in Article Google Scholar

    [84] Yan W., Hu J., Zeng H., et al. (2025). The application of large language models in primary healthcare services and the challenges. Chinese General Practice 28:1−6. DOI:10.12114/j.issn.1007-9572.2024.0277

    View in Article CrossRef Google Scholar

    [85] Tsinghua University Vanke School of Public Health, Peking University School of Public Health, Chinese Association of General Practitioners of Chinese Medical Doctor Association. (2025). Chinese expert consensus on artificial intelligence general practitioner (AIGP). Chinese General Practice 28:135−142. DOI:10.12114/j.issn.1007-9572.2024.0453

    View in Article CrossRef Google Scholar

    [86] Liu J. and Liang W. (2024). Technological innovations to address health challenges. Chinese General Practice 27:3449−3452. DOI:10.12114/j.issn.1007-9572.2024.0210

    View in Article CrossRef Google Scholar

    [87] Yuan B., Dai H., Wu J., et al. (2021). Application of artificial intelligence applications in general practice. Chinese General Practice 19:1433−1436,1572. DOI:10.16766/j.cnki.issn.1674-4152.002079

    View in Article CrossRef Google Scholar

    [88] Li Y. H., Li Y. L., Wei M. Y., et al. (2024). Innovation and challenges of artificial intelligence technology in personalized healthcare. Sci. Rep. 14:18994. DOI:10.1038/s41598-024-70073-7

    View in Article CrossRef Google Scholar

    [89] Lopez R. P., Laleh N. G., Mahmood F., et al. (2024). A guide to artificial intelligence for cancer researchers. Nat. Rev. Cancer 24:427−441. DOI:10.1038/s41568-024-00694-7

    View in Article CrossRef Google Scholar

    [90] Michael M., Oishi B., Hossein A. Z. S., et al. (2023). Foundation models for generalist medical artificial intelligence. Nature 616:259−265. DOI:10.1038/s41586-023-05881-4

    View in Article CrossRef Google Scholar

    [91] Shi Y. (2024). Drug development in the AI era: AlphaFold 3 is coming! The Innovation 5:100685. DOI:10.1016/j.xinn.2024.100685

    View in Article Google Scholar

    [92] Xu Y., Liu X., Cao X., et al. (2021). Artificial intelligence: A powerful paradigm for scientific research. The Innovation 2:100179. DOI:10.1016/j.xinn.2021.100179

    View in Article CrossRef Google Scholar

    [93] Huang T. and Li Y. (2023). Current progress, challenges, and future perspectives of language models for protein representation and protein design. The Innovation 4:100446. DOI:10.1016/j.xinn.2023.100446

    View in Article CrossRef Google Scholar

    [94] Freifeld C. C., Mandl K. D., Reis B. Y., et al. (2008). HealthMap: Global infectious disease monitoring through automated classification and visualization of Internet media reports. J. Am. Med. Inform. Assoc. 15:150−157. DOI:10.1197/jamia.M2544

    View in Article CrossRef Google Scholar

    [95] Baddal B., Taner F. and Uzun Ozsahin D. (2024). Harnessing of artificial intelligence for the diagnosis and prevention of hospital-acquired infections: A systematic review. Diagnostics 14:484. DOI:10.3390/diagnostics14050484

    View in Article CrossRef Google Scholar

    [96] Zhou Z., Yang T. and Hu K. (2023). Traditional chinese medicine epidemic prevention and treatment question-answering model based on llms. 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) pp:4755-4760. DOI:10.1109/BIBM58861.2023.10385748

    View in Article Google Scholar

    [97] Rosa G. J. (2010). The Elements of Statistical Learning. Oxford University Press. https://www.sas.upenn.edu/~fdiebold/NoHesitations/BookAdvanced.pdf.

    View in Article Google Scholar

    [98] Ribeiro M. T., Singh S. and Guestrin C. (2016). "Why Should I Trust You?": Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining pp:1135-1144. DOI:10.48550/arXiv.1602.04938

    View in Article Google Scholar

    [99] Shrikumar A., Greenside P. and Kundaje A. (2017). Learning important features through propagating activation differences. International conference on machine learning pp:3145-3153. DOI:10.48550/arXiv.1704.02685

    View in Article Google Scholar

    [100] Lundberg S. and Lee S. (2017). A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874. DOI:10.48550/arXiv.1705.07874

    View in Article Google Scholar

    [101] He X. and Lin X. (2020). Challenges and opportunities in statistics and data science: Ten research areas. Harv. Data Sci. Rev. 2:10.1162/99608f92.95388fcb. DOI:10.1162/99608f92.95388fcb

    View in Article Google Scholar

    [102] Yang Z., Zhang A. and Sudjianto A. (2020). GAMI-net: An explainable neural network based on generalized additive models with structured interactions. arXiv preprint arXiv:2003.07132. DOI:10.48550/arXiv.2003.07132

    View in Article Google Scholar

    [103] Agarwal R., Frosst N., Zhang X., et al. (2020). Neural additive models: Interpretable machine learning with neural nets. arXiv preprint arXiv:2004.13912. DOI:10.48550/arXiv.2004.13912

    View in Article Google Scholar

    [104] Aubry M. and Russell B. C. (2015). Understanding deep features with computer-generated imagery. 2015 IEEE International Conference on Computer Vision (ICCV):2875-2883. DOI:10.1109/iccv.2015.329

    View in Article Google Scholar

    [105] Rauber P. E., Fadel S. G., Falcao A. X., et al. (2017). Visualizing the hidden activity of artificial neural networks. IEEE Transactions on Visualization and Computer Graphics 23:101−110. DOI:10.1109/tvcg.2016.2598838

    View in Article CrossRef Google Scholar

    [106] Hong S. and Fan X. (2024). Reconciling explanations in multi-model systems through probabilistic argumentation. arXiv preprint arXiv:2404.13419. DOI:10.48550/arXiv.2404.13419

    View in Article Google Scholar

    [107] Goldstein B. A., Navar A. M. and Carter R. E. (2017). Moving beyond regression techniques in cardiovascular risk prediction: applying machine learning to address analytic challenges. Eur. Heart J. 38:1805−1814. DOI:10.1093/eurheartj/ehw302

    View in Article CrossRef Google Scholar

    [108] Bzdok D., Engemann D. and Thirion B. (2020). Inference and prediction diverge in biomedicine. Patterns 1:100119. DOI:10.1016/j.patter.2020.100119

    View in Article CrossRef Google Scholar

    [109] Peters J., Janzing D. and Schlkopf B. (2017). Elements of Causal Inference: Foundations and Learning Algorithms (The MIT Press). https://library.oapen.org/bitstream/id/056a11be-ce3a-44b9-8987-a6c68fce8d9b/11283.pdf.

    View in Article Google Scholar

    [110] Blakely T., Lynch J., Simons K., et al. (2019). Reflection on modern methods: When worlds collide-prediction, machine learning and causal inference. Int. J. Epidemiol. 49:2058−2064. DOI:10.1093/ije/dyz132

    View in Article CrossRef Google Scholar

    [111] Sanchez P., Voisey J., Xia T., et al. (2022). Causal machine learning for healthcare and precision medicine. R. Soc. Open Sci. 9:220638. DOI:10.1098/rsos.220638

    View in Article CrossRef Google Scholar

    [112] Yao M., Miller G. W., Vardarajan B. N., et al. Deciphering proteins in Alzheimer's disease: A new Mendelian randomization method integrated with AlphaFold3 for 3D structure prediction. Cell Genom. 4:100700. DOI:10.1016/j.xgen.2024.100700

    View in Article Google Scholar

    [113] So A., Hooshyar D., Park K. W., et al. (2017). Early diagnosis of dementia from clinical data by machine learning techniques. Appl. Sci. 7:651. DOI:10.3390/app7070651

    View in Article CrossRef Google Scholar

    [114] Parafita Á. and Vitrià J. (2021). Deep Causal Graphs for Causal Inference, Black-Box Explainability and Fairness. International Conference of the Catalan Association for Artificial Intelligence. DOI:10.3233/FAIA210162

    View in Article Google Scholar

    [115] Zapata D., Meyer M. and Muller O. (2024). Bridging the Gap Between Data-Driven and Theory-Driven Modelling -- Leveraging Causal Machine Learning for Integrative Modelling of Dynamical Systems. https://synthical.com/article/Bridging-the-Gap-Between-Data-Driven-and-Theory-Driven-Modelling--Leveraging-Causal-Machine-Learning-for-Integrative-Modelling-of-Dynamical-Systems-166a4e01-b49d-4574-8302-e217d3e706bc.

    View in Article Google Scholar

    [116] Giudice E., Kuipers J. and Moffa G. (2024). Bayesian causal inference with gaussian process networks. arXiv preprint arXiv:2402.00623. DOI:10.48550/arXiv.2402.00623

    View in Article Google Scholar

    [117] Wang B., Li J., Chang H., et al. (2025). Heterophilic graph neural networks optimization with causal message-passing. Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining pp:829-837. DOI:10.1145/3701551.3703568

    View in Article Google Scholar

    [118] Gimenez J. R. and Rothenhäusler D. (2021). Causal aggregation: Estimation and inference of causal effects by constraint-based data fusion. J. Mach. Learn. Res. 23:335:331-335:360. https://www.jmlr.org/papers/volume23/21-0656/21-0656.pdf.

    View in Article Google Scholar

    [119] Ma C. and Zhang C. (2023). High precision causal model evaluation with conditional randomization. arXiv preprint arXiv:2311.01902. DOI:10.48550/arXiv.2311.01902

    View in Article Google Scholar

    [120] Yadav P., Pruinelli L., Hoff A., et al. (2016). Causal inference in observational data. arXiv preprint arXiv:1611.04660. DOI:10.48550/arXiv.1611.04660

    View in Article Google Scholar

    [121] Wiegrebe S., Kopper P., Sonabend R., et al. (2024). Deep learning for survival analysis: A review. Artif. Intell. Rev. 57:65. DOI:10.1007/s10462-023-10681-3

    View in Article CrossRef Google Scholar

    [122] Olden J. D. and Jackson D. A. (2002). Illuminating the “black box”: A randomization approach for understanding variable contributions in artificial neural networks. Ecol. modell. 154:135−150. DOI:10.1016/s0304-3800(02)00064-9

    View in Article CrossRef Google Scholar

    [123] Refenes A. P. and Zapranis A. (1999). Neural model identification, variable selection and model adequacy. J. Forecast. 18:299−332. DOI:3.0.co;2-t">10.1002/(sici)1099-131x(199909)18:5<299::aid-for725>3.0.co;2-t

    View in Article CrossRef Google Scholar

    [124] Fan Y. and Li Q. (1996). Consistent model specification tests: omitted variables and semiparametric functional forms. Econometrica 64:865−890. DOI:10.2307/2171848

    View in Article CrossRef Google Scholar

    [125] Lavergne P. and Vuong Q. (2000). Nonparametric significance testing. Econom. Theory 16:576−601. DOI:10.1017/s0266466600164059

    View in Article CrossRef Google Scholar

    [126] Gu J., Li D. and Liu D. (2007). Bootstrap non-parametric significance test. J. Nonparametr. Stat. 19:215−230. DOI:10.1080/10485250701734497

    View in Article CrossRef Google Scholar

    [127] Horel E. and Giesecke K. (2020). Significance tests for neural networks. JMLR 21:1−29. DOI:10.48550/arXiv.1902.06021

    View in Article CrossRef Google Scholar

    [128] Liu Z., Li Z., Wang J., et al. (2024). Full Bayesian Significance Testing for Neural Networks. Proceedings of the AAAI Conference on Artificial Intelligence 38:8841−8849. DOI:10.1609/aaai.v38i8.28731

    View in Article CrossRef Google Scholar

    [129] McKinney S. M., Sieniek M., Godbole V., et al. (2020). International evaluation of an AI system for breast cancer screening. Nature 577:89−94. DOI:10.1038/s41586-019-1799-6

    View in Article CrossRef Google Scholar

    [130] Yang Y.-C. and Chen H.-H. (2025). Dynamic DropConnect: Enhancing neural network robustness through adaptive edge dropping strategies. arXiv preprint arXiv:2502.19948. DOI:10.48550/arXiv.2502.19948

    View in Article Google Scholar

    [131] Tack J., Yu S., Jeong J., et al. (2022). Consistency Regularization for Adversarial Robustness. Proceedings of the AAAI conference on artificial intelligence 36:8414−8422. DOI:10.1609/aaai.v36i8.20817

    View in Article CrossRef Google Scholar

    [132] Chen R. J., Wang J. J., Williamson D. F. K., et al. (2023). Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng. 7:719−742. DOI:10.1038/s41551-023-01056-8

    View in Article CrossRef Google Scholar

    [133] Liu M., Ning Y., Teixayavong S., et al. (2023). A translational perspective towards clinical AI fairness. NPJ Digit. Med. 6:172. DOI:10.1038/s41746-023-00918-4

    View in Article CrossRef Google Scholar

    [134] Obermeyer Z., Powers B., Vogeli C., et al. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366:447−453. DOI. DOI:10.1126/science.aax2342

    View in Article CrossRef Google Scholar

    [135] Hasanzadeh F., Josephson C. B., Waters G., et al. (2025). Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit. Med. 8:154. DOI:10.1038/s41746-025-01503-7

    View in Article CrossRef Google Scholar

    [136] Taskesen B., Blanchet J., Kuhn D., et al. (2021). A Statistical Test for Probabilistic Fairness. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency pp:648–665. DOI:10.1145/3442188.3445927

    View in Article Google Scholar

    [137] Jin Y. (2023). Robust Statistical Inference with Privacy and Fairness Constraints. UC Davis. https://escholarship.org/uc/item/4g91c7rr.

    View in Article Google Scholar

    [138] Zhao Y., Wong Z. S.-Y. and Tsui K. L. (2018). A framework of rebalancing imbalanced healthcare data for rare events’ classification: A case of look-alike sound-alike mix-up incident detection. J. Healthcare Eng. 2018:1−11. DOI:10.1155/2018/6275435

    View in Article CrossRef Google Scholar

    [139] Gan D.-S. H. Y., Bevilacqua V. and Figueroa J. C. (2005). Advanced Intelligent Computing Technology and Applications. Proceedings of the 2024 International Conference on Intelligent Computing DOI:10.1007/978-981-97-5612-4

    View in Article Google Scholar

    [140] He H., Bai Y., Garcia E. A., et al. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. Proceedings of the 2008 IEEE international joint conference on neural networks (IEEE world congress on computational intelligence) pp:1322-1328. DOI:10.1109/IJCNN.2008.4633969

    View in Article Google Scholar

    [141] Mathew J., Luo M., Pang C. K., et al. (2015). Kernel-based SMOTE for SVM classification of imbalanced datasets. Proceedings of the IECON 2015-41ST annual conference of the IEEE industrial electronics society pp:001127-001132. DOI:10.1109/IECON.2015.7392251

    View in Article Google Scholar

    [142] Wang J., Wang K., Yu Y., et al. (2025). Self-improving generative foundation model for synthetic medical image generation and clinical applications. Nat. Med. 31:609−617. DOI:10.1038/s41591-024-03359-y

    View in Article CrossRef Google Scholar

    [143] Pan X., Cheng J., Hou F., et al. (2023). SMILE: Cost-sensitive multi-task learning for nuclear segmentation and classification with imbalanced annotations. Med. Image Anal. 88:102867. DOI:10.1016/j.media.2023.102867

    View in Article CrossRef Google Scholar

    [144] Yeung M., Sala E., Schönlieb C.-B., et al. (2022). Unified Focal loss: Generalising Dice and cross entropy-based losses to handle class imbalanced medical image segmentation. Comput. Med. Imaging Graph. 95:102026. DOI:10.1016/j.compmedimag.2021.102026

    View in Article CrossRef Google Scholar

    [145] Wang S., Zhou T., Shen Y., et al. (2025). Generative AI Enables EEG Super-Resolution via Spatio-Temporal Adaptive Diffusion Learning. IEEE Transactions on Consumer Electronics 71:1034-1045. https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=10839074.

    View in Article Google Scholar

    [146] Li Y., Wang Y., Lei B., et al. (2025). SCDM: Unified representation learning for EEG-to-fNIRS cross-modal generation in MI-BCIs. IEEE Trans. Med. Imaging 44:2384−2394. DOI:10.1109/TMI.2025.3532480

    View in Article CrossRef Google Scholar

    [147] Colasacco C. J. and Born H. L. (2024). A case of artificial intelligence chatbot hallucination. JAMA Otolaryngol. Head Neck Surg. 150:457−458. DOI:10.1001/jamaoto.2024.0428

    View in Article CrossRef Google Scholar

    [148] Arif T. B., Munaf U. and Ul-Haque I. (2023). The future of medical education and research: Is ChatGPT a blessing or blight in disguise. Med. Educ. Online 28:2181052. DOI:10.1080/10872981.2023.2181052

    View in Article CrossRef Google Scholar

    [149] Gordijn B. and Have H. T. (2023). ChatGPT: Evolution or revolution. Med. Health Care Philos. 26:1−2. DOI:10.1007/s11019-023-10136-0

    View in Article CrossRef Google Scholar

    [150] Howard A., Hope W. and Gerada A. (2023). ChatGPT and antimicrobial advice: The end of the consulting infection doctor. Lancet Infect. Dis. 23:405−406. DOI:10.1016/S1473-3099(23)00113-5

    View in Article CrossRef Google Scholar

    [151] Kelly C. J., Karthikesalingam A., Suleyman M., et al. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17:195. DOI:10.1186/s12916-019-1426-2

    View in Article CrossRef Google Scholar

    [152] Aljamaan F., Temsah M.-H., Altamimi I., et al. (2024). Reference hallucination score for medical artificial intelligence chatbots: Development and usability study. JMIR Med. Inform. 12:e54345. DOI:10.2196/54345

    View in Article CrossRef Google Scholar

    [153] Su J., Kempe J. and Ullrich K. (2024). Mission impossible: A statistical perspective on jailbreaking LLMs. arXiv preprint arXiv:2408.01420. DOI:10.48550/arXiv.2408.01420

    View in Article Google Scholar

    [154] W. H. O. (2010). Monitoring the building blocks of health systems: A handbook of indicators and their measurement strategies (World Health Organization). https://iris.who.int/bitstream/handle/10665/258734/9789241564052-eng.pdf.

    View in Article Google Scholar

    [155] Chan M., Kazatchkine M., Lob-Levyt J., et al. (2010). Meeting the demand for results and accountability: A call for action on health data from eight global health agencies. PLoS Med. 7:e1000223. DOI:10.1371/journal.pmed.1000223

    View in Article CrossRef Google Scholar

    [156] Chinta S. V., Wang Z., Zhang X., et al. (2024). AI-driven healthcare: A survey on ensuring fairness and mitigating bias. arXiv preprint arXiv:2407.19655. DOI:10.48550/arXiv.2407.19655

    View in Article Google Scholar

    [157] Rudroff T., Rainio O. and Klén R. (2024). AI for the prediction of early stages of Alzheimer's disease from neuroimaging biomarkers–A narrative review of a growing field. Neurol. Sci. 45:5117−5127. DOI:10.1007/s10072-024-07649-8

    View in Article CrossRef Google Scholar

    [158] Patharkar A., Cai F., Al-Hindawi F., et al. (2024). Predictive modeling of biomedical temporal data in healthcare applications: Review and future directions. Front. Physiol. 15:1386760. DOI:10.3389/fphys.2024.1386760

    View in Article CrossRef Google Scholar

    [159] Janga B., Asamani G. P., Sun Z., et al. (2023). A review of practical ai for remote sensing in earth sciences. Remote Sens. 15:4112. DOI:10.3390/rs15164112

    View in Article CrossRef Google Scholar

    [160] Chen H., Hailey D., Wang N., et al. (2014). A review of data quality assessment methods for public health information systems. International journal of environmental research and public health 11:5170−5207. DOI:10.48550/arXiv.2502.19948

    View in Article CrossRef Google Scholar

    [161] Vollmer S., Mateen B. A., Bohner G., et al. (2020). Machine learning and artificial intelligence research for patient benefit: 20 critical questions on transparency, replicability, ethics, and effectiveness. BMJ 368:l6927. DOI:10.1136/bmj.l6927

    View in Article CrossRef Google Scholar

    [162] Haibe-Kains B., Adam G. A., Hosny A., et al. (2020). Transparency and reproducibility in artificial intelligence. Nature 586:E14−E16. DOI:10.1038/s41586-020-2766-y

    View in Article CrossRef Google Scholar

    [163] Pearl J. (2009). Causality: Models, Reasoning and Inference (Cambridge University Press). Econometric Theory 19:675−685. DOI:10.1017/S0266466603004109

    View in Article CrossRef Google Scholar

    [164] Andrieu C., De Freitas N., Doucet A., et al. (2003). An introduction to MCMC for machine learning. Mach. Learn. 50:5−43. DOI:10.1023/A:1020281327116

    View in Article CrossRef Google Scholar

    [165] Hunter D. J. and Holmes C. (2023). Where medical statistics meets artificial intelligence. N. Engl. J. Med. 389:1211−1219. DOI:10.1056/NEJMra2212850

    View in Article CrossRef Google Scholar

    [166] Bragazzi N. L. and Garbarino S. (2024). Toward clinical generative AI: Conceptual framework. JMIR AI 3:e55957. DOI:10.2196/55957

    View in Article CrossRef Google Scholar

    [167] Liu L., Yang X., Lei J., et al. (2024). A survey on medical large language models: Technology, application, trustworthiness, and future directions. arXiv preprint arXiv:2406.03712. DOI:10.48550/arXiv.2406.03712

    View in Article Google Scholar

    [168] He K., Mao R., Lin Q., et al. (2025). A survey of large language models for healthcare: From data, technology, and applications to accountability and ethics. Inform. Fusion 118:102963. DOI:10.1016/j.inffus.2025.102963

    View in Article CrossRef Google Scholar

    [169] Wang D. and Zhang S. (2024). Large language models in medical and healthcare fields: applications, advances, and challenges. Artif. Intell. Rev. 57:299. DOI:10.1007/s10462-024-10921-0

    View in Article CrossRef Google Scholar

    [170] Joshi G., Jain A., Araveeti S. R., et al. (2024). FDA-approved artificial intelligence and machine learning (AI/ML)-enabled medical devices: An updated landscape. Electronics 13:498. DOI:10.3390/electronics13030498

    View in Article CrossRef Google Scholar

    [171] Schaffter T., Buist D. S., Lee C. I., et al. (2020). Evaluation of combined artificial intelligence and radiologist assessment to interpret screening mammograms. JAMA Netw. Open 3:e200265−200265. DOI:10.1001/jamanetworkopen.2020.0265

    View in Article CrossRef Google Scholar

    [172] Houssami N., Kirkpatrick-Jones G., Noguchi N., et al. (2019). Artificial Intelligence (AI) for the early detection of breast cancer: A scoping review to assess AI’s potential in breast screening practice. Expert Rev. Med. Devices 16:351−362. DOI:10.1080/17434440.2019.1610387

    View in Article CrossRef Google Scholar

    [173] Park S. H. and Han K. (2018). Methodologic guide for evaluating clinical performance and effect of artificial intelligence technology for medical diagnosis and prediction. Radiology 286:800−809. DOI:10.1148/radiol.2017171920

    View in Article CrossRef Google Scholar

    [174] Ouyang D. and Hogan J. (2024). We need more randomized clinical trials of AI. NEJM AI 1:AIe2400881. DOI:10.1056/AIe2400881

    View in Article CrossRef Google Scholar

    [175] Sedano R., Solitano V., Vuyyuru S. K., et al. (2025). Artificial intelligence to revolutionize IBD clinical trials: A comprehensive review. Therap. Adv. Gastroenterol. 18:17562848251321915. DOI:10.1177/17562848251321915

    View in Article Google Scholar

  • Cite this article:

    Wei Q., Cui M., Liu Z., et al. (2025). Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI. The Innovation Medicine 3:100145. https://doi.org/10.59717/j.xinn-med.2025.100145
    Wei Q., Cui M., Liu Z., et al. (2025). Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI. The Innovation Medicine 3:100145. https://doi.org/10.59717/j.xinn-med.2025.100145

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(2)     Tables(2)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(12102) PDF downloads(2262)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint