MACHINE LEARNING IN PREDICTION OF SEPSIS IN INTENSIVE CARE UNITS: A LITERATURE REVIEW OF ALGORITHMIC MODELS AND CLINICAL IMPLEMENTATION
DOI:
https://doi.org/10.31435/ijitss.3(51).2026.5768Keywords:
Sepsis; Machine Learning; Intensive Care; TREWS; Epic Sepsis Model; COMPOSER; Sepsis ImmunoScore; PRISMAAbstract
Background. Sepsis continues to be among the leading causes of in-hospital death, accounting for around 11 million deaths globally in 2017 and pooled ICU mortality of nearly 42% (Rudd et al., 2020; Fleischmann-Struzek et al., 2020). The chance of survival decreases with each additional hour of delay in starting effective antibiotic treatment (Kumar et al., 2006), and standard screening tools, including SIRS and qSOFA, are known to overlook a meaningful share of cases (Raith et al., 2017). Models trained on routinely captured EHR data have therefore been put forward as a means of recognising patients at risk well before overt deterioration.
Objective. To bring together the peer-reviewed evidence published between 2015 and 2025 on machine-learning algorithms for predicting sepsis in adult ICU patients, with attention to model architecture, prospective validation and the regulatory status of systems that have actually reached the bedside.
Methods. A PRISMA 2020-aligned narrative review (Page et al., 2021) covering MEDLINE, Scopus, Web of Science, Embase and IEEE Xplore from January 2015 to March 2025. Reporting quality was checked against TRIPOD+AI (Collins et al., 2024) and risk of bias with PROBAST (Wolff et al., 2019). Seventy-three studies met the inclusion criteria.
Results. Gradient-boosted trees, LSTM networks and transformer-based models reach internal AUROCs of about 0.83 to 0.94, while qSOFA and SIRS sit between 0.69 and 0.76 (Henry et al., 2015; Nemati et al., 2018). TREWS was associated with an 18.7% relative drop in adjusted in-hospital sepsis mortality (Adams et al., 2022); the only published randomised trial of an ML sepsis predictor, InSight, lowered mortality from 21.3% to 8.96% in a small ICU cohort (Shimabukuro et al., 2017); COMPOSER showed a 17% relative reduction at the bedside (Boussina et al., 2024). By contrast, the widely deployed Epic Sepsis Model achieved an external AUROC of only 0.63 (Wong et al., 2021). Two AI sepsis diagnostics have so far cleared the FDA: Sepsis ImmunoScore (De Novo DEN230036, 2024) and TriVerity (510(k) K241676, 2025).
Conclusions. On the whole ML models discriminate sepsis better than traditional scores, but uneven outcome definitions, training-set bias, alert fatigue and patchy external validation still limit how far the gains travel between centres. Prospective trials reported to TRIPOD+AI standards, together with explicit equity audits, will be needed before ML sepsis prediction can be treated as routine care.
References
Adams, R., Henry, K. E., Sridharan, A., Soleimani, H., Zhan, A., Rawat, N., Johnson, L., Hager, D. N., Cosgrove, S. E., Markowski, A., Klein, E. Y., Chen, E. S., Saheed, M. O., Henley, M., Miranda, S., Houston, K., Linton, R. C., Ahluwalia, A. R., Wu, A. W., & Saria, S. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455–1460. https://doi.org/10.1038/s41591-022-01894-0
Andaur Navarro, C. L., Damen, J. A. A., Takada, T., Nijman, S. W. J., Dhiman, P., Ma, J., Collins, G. S., Bajpai, R., Riley, R. D., Moons, K. G. M., & Hooft, L. (2021). Risk of bias in studies on prediction models developed using supervised machine learning techniques: Systematic review. BMJ, 375, Article n2281. https://doi.org/10.1136/bmj.n2281
Angus, D. C., Linde-Zwirble, W. T., Lidicker, J., Clermont, G., Carcillo, J., & Pinsky, M. R. (2001). Epidemiology of severe sepsis in the United States. Critical Care Medicine, 29(7), 1303–1310.
Bhargava, A., Lopez-Espina, C., Schmalz, L., Khan, S., Watson, G. L., Urdiales, D., Updike, L., Kurtzman, N. D., Dhamija, A., Liu, A., Berryman, F., Davis, R. A., Hurst, J., Taylor, J., Mann, D. K., Babu, B. G., Gupta, R., Alsakka, Z., Ghita, C., . . . Ross, J. J. (2024). FDA-authorized AI/ML tool for sepsis prediction: Development and validation. NEJM AI, 1(12). https://doi.org/10.1056/AIoa2400867
Boussina, A., Shashikumar, S. P., Malhotra, A., Owens, R. L., El-Kareh, R., Longhurst, C. A., Quintero, K., Donahue, A., Chan, T. C., Nemati, S., & Wardi, G. (2024). Impact of a deep learning sepsis prediction model on quality of care and survival. npj Digital Medicine, 7, Article 14. https://doi.org/10.1038/s41746-023-00986-6
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.
Buchman, T. G., Simpson, S. Q., Sciarretta, K. L., et al. (2020). Sepsis among Medicare beneficiaries: 1. The burdens of sepsis, 2012–2018. Critical Care Medicine, 48(3), 276–288.
Calvert, J. S., Price, D. A., Chettipally, U. K., et al. (2016). A computational approach to early sepsis detection. Computers in Biology and Medicine, 74, 69–73.
Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in health care. Annual Review of Biomedical Data Science, 4, 123–144.
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794).
Collins, G. S., Moons, K. G. M., Dhiman, P., et al. (2024). TRIPOD+AI statement. BMJ, 385, Article e078378.
Cruz Rivera, S., Liu, X., Chan, A.-W., Denniston, A. K., & Calvert, M. J. (2020). SPIRIT-AI extension. Nature Medicine, 26, 1351–1363.
Davis, S. E., Lasko, T. A., Chen, G., Siew, E. D., & Matheny, M. E. (2017). Calibration drift in regression and machine learning models for acute kidney injury. Journal of the American Medical Informatics Association, 24(6), 1052–1061.
Desautels, T., Calvert, J., Hoffman, J., et al. (2016). Prediction of sepsis in the ICU with minimal EHR data. JMIR Medical Informatics, 4(3), Article e28.
Epic Systems. (2018). Sepsis predictive analytics model. Epic Systems Corporation.
European Union. (2017). Regulation (EU) 2017/745 on medical devices (MDR). Official Journal of the European Union, L 117.
European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
Evans, L., Rhodes, A., Alhazzani, W., et al. (2021). Surviving Sepsis Campaign: International guidelines for management of sepsis and septic shock 2021. Critical Care Medicine, 49(11), e1063–e1143.
Faltys, M., Zimmermann, M., Lyu, X., et al. (2021). HiRID, a high time-resolution ICU dataset (Version 1.1.1) [Data set]. PhysioNet.
Finlayson, S. G., Subbaswamy, A., Singh, K., et al. (2021). The clinician and dataset shift in artificial intelligence. The New England Journal of Medicine, 385(3), 283–286.
Fleischmann-Struzek, C., Mellhammar, L., Rose, N., et al. (2020). Incidence and mortality of hospital- and ICU-treated sepsis. Intensive Care Medicine, 46(8), 1552–1562.
Fleuren, L. M., Klausch, T. L. T., Zwager, C. L., et al. (2020). Machine learning for the prediction of sepsis: A systematic review and meta-analysis. Intensive Care Medicine, 46(3), 383–400.
Freund, Y., Lemachatti, N., Krastinova, E., et al. (2017). Prognostic accuracy of Sepsis-3 criteria for in-hospital mortality. JAMA, 317(3), 301–308.
Futoma, J., Hariharan, S., & Heller, K. (2017). Learning to detect sepsis with a multitask Gaussian process RNN classifier. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1174–1182).
Futoma, J., Simons, M., Panch, T., Doshi-Velez, F., & Celi, L. A. (2020). The myth of generalisability in clinical research and machine learning in health care. The Lancet Digital Health, 2(9), e489–e492.
Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745–e750.
Goh, K. H., Wang, L., Yeow, A. Y. K., et al. (2021). Artificial intelligence in sepsis early prediction. Nature Communications, 12, Article 711.
Gottesman, O., Johansson, F., Komorowski, M., et al. (2019). Guidelines for reinforcement learning in healthcare. Nature Medicine, 25(1), 16–18.
Habib, A. R., Lin, A. L., & Grant, R. W. (2021). The Epic Sepsis Model falls short. JAMA Internal Medicine, 181(8), 1040–1041.
Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems, 29, 3323–3331.
Henry, K. E., Hager, D. N., Pronovost, P. J., & Saria, S. (2015). A targeted real-time early warning score (TREWScore) for septic shock. Science Translational Medicine, 7(299), Article 299ra122.
Henry, K. E., Adams, R., Parent, C., et al. (2022). Factors driving provider adoption of the TREWS system. Nature Medicine, 28(7), 1447–1454.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780.
Jeter, R., Josef, C., Shashikumar, S., & Nemati, S. (2019). Does the “AI Clinician” learn optimal treatment strategies for sepsis in intensive care? [Preprint]. arXiv. https://arxiv.org/abs/1902.03271
Johnson, A. E. W., Pollard, T. J., Shen, L., et al. (2016). MIMIC-III, a freely accessible critical care database. Scientific Data, 3, Article 160035.
Johnson, A. E. W., Bulgarelli, L., Shen, L., et al. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10, Article 1.
Jones, B. E., Collingridge, D. S., Vines, C. G., Post, H., Holmen, J., Allen, T. L., Haug, P. J., Weir, C. R., & Dean, N. C. (2019). Clinical decision support in a learning health care system: Identifying physicians’ reasons for rejection of best-practice recommendations in pneumonia through computerized clinical decision support. Applied Clinical Informatics, 10(1), 1–9.
Kam, H. J., & Kim, H. Y. (2017). Learning representations for the early detection of sepsis with deep neural networks. Computers in Biology and Medicine, 89, 248–255.
Kaukonen, K.-M., Bailey, M., Pilcher, D., Cooper, D. J., & Bellomo, R. (2015). Systemic inflammatory response syndrome criteria in defining severe sepsis. The New England Journal of Medicine, 372(17), 1629–1638.
Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., & Faisal, A. A. (2018). The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine, 24(11), 1716–1720.
Kumar, A., Roberts, D., Wood, K. E., et al. (2006). Duration of hypotension before initiation of effective antimicrobial therapy. Critical Care Medicine, 34(6), 1589–1596.
Lauritsen, S. M., Kalør, M. E., Kongsgaard, E. L., et al. (2020). Early detection of sepsis utilizing deep learning on electronic health record event sequences. Artificial Intelligence in Medicine, 104, Article 101820.
Liu, V. X., Fielding-Singh, V., Greene, J. D., et al. (2017). The timing of early antibiotics and hospital mortality in sepsis. American Journal of Respiratory and Critical Care Medicine, 196(7), 856–863.
Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., & Denniston, A. K. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nature Medicine, 26(9), 1364–1374.
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.
Lundberg, S. M., Nair, B., Vavilala, M. S., et al. (2018). Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nature Biomedical Engineering, 2(10), 749–760.
Mayhew, M. B., Buturovic, L., Luethy, R., et al. (2020). A generalizable 29-mRNA neural-network classifier for acute bacterial and viral infections. Nature Communications, 11, Article 1177.
Mayr, F. B., Yende, S., Linde-Zwirble, W. T., et al. (2010). Infection rate and acute organ dysfunction risk as explanations for racial differences in severe sepsis. JAMA, 303(24), 2495–2503.
Moor, M., Rieck, B., Horn, M., Jutzeler, C. R., & Borgwardt, K. (2021). Early prediction of sepsis in the ICU using machine learning: A systematic review. Frontiers in Medicine, 8, Article 607952.
Nemati, S., Holder, A., Razmi, F., et al. (2018). An interpretable machine learning model for accurate prediction of sepsis in the ICU. Critical Care Medicine, 46(4), 547–553.
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.
Ouzzani, M., Hammady, H., Fedorowicz, Z., & Elmagarmid, A. (2016). Rayyan—A web and mobile app for systematic reviews. Systematic Reviews, 5, Article 210.
Page, M. J., McKenzie, J. E., Bossuyt, P. M., et al. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, Article n71.
Pollard, T. J., Johnson, A. E. W., Raffa, J. D., et al. (2018). The eICU Collaborative Research Database, a freely available multicenter database for critical care research. Scientific Data, 5, Article 180178.
Raith, E. P., Udy, A. A., Bailey, M., et al. (2017). Prognostic accuracy of the SOFA score, SIRS criteria, and qSOFA score for in-hospital mortality. JAMA, 317(3), 290–300.
Reyna, M. A., Josef, C. S., Jeter, R., et al. (2020). Early prediction of sepsis from clinical data: The PhysioNet/Computing in Cardiology Challenge 2019. Critical Care Medicine, 48(2), 210–217.
Rhee, C., Dantes, R., Epstein, L., et al. (2017). Incidence and trends of sepsis in US hospitals using clinical vs claims data, 2009–2014. JAMA, 318(13), 1241–1249.
Rhee, C., Jones, T. M., Hamad, Y., et al. (2019). Prevalence, underlying causes, and preventability of sepsis-associated mortality in US acute care hospitals. JAMA Network Open, 2(2), Article e187571.
Rieke, N., Hancox, J., Li, W., et al. (2020). The future of digital health with federated learning. npj Digital Medicine, 3, Article 119.
Rudd, K. E., Johnson, S. C., Agesa, K. M., et al. (2020). Global, regional, and national sepsis incidence and mortality, 1990–2017: Analysis for the Global Burden of Disease Study. The Lancet, 395(10219), 200–211.
Rudin, C. (2019). Stop explaining black box machine learning models for high-stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215.
Scherpf, M., Grasser, F., Malberg, H., & Zaunseder, S. (2019). Predicting sepsis with a recurrent neural network using the MIMIC III database. Computers in Biology and Medicine, 113, Article 103395.
Sendak, M. P., Ratliff, W., Sarro, D., et al. (2020). Real-world integration of a sepsis deep learning technology into routine clinical care: Implementation study. JMIR Medical Informatics, 8(7), Article e15182.
Seymour, C. W., Gesten, F., Prescott, H. C., et al. (2017). Time to treatment and mortality during mandated emergency care for sepsis. The New England Journal of Medicine, 376(23), 2235–2244.
Seymour, C. W., Kennedy, J. N., Wang, S., et al. (2019). Derivation, validation, and potential treatment implications of novel clinical phenotypes for sepsis. JAMA, 321(20), 2003–2017.
Shashikumar, S. P., Josef, C. S., Sharma, A., & Nemati, S. (2021). DeepAISE: A recurrent neural survival model for early prediction of sepsis. Artificial Intelligence in Medicine, 113, Article 102036.
Shimabukuro, D. W., Barton, C. W., Feldman, M. D., Mataraso, S. J., & Das, R. (2017). Effect of a machine learning-based severe sepsis prediction algorithm on patient survival and hospital length of stay: A randomised clinical trial. BMJ Open Respiratory Research, 4(1), Article e000234.
Singer, M., Deutschman, C. S., Seymour, C. W., et al. (2016). The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA, 315(8), 801–810.
Singhal, K., Azizi, S., Tu, T., et al. (2023). Large language models encode clinical knowledge. Nature, 620, 172–180.
Sterne, J. A. C., Savović, J., Page, M. J., et al. (2019). RoB 2: A revised tool for assessing risk of bias in randomised trials. BMJ, 366, Article l4898.
Subbaswamy, A., & Saria, S. (2020). From development to deployment: Dataset shift, causality, and shift-stable models in health AI. Biostatistics, 21(2), 345–352.
Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 3319–3328).
Thoral, P. J., Peppink, J. M., Driessen, R. H., et al. (2021). Sharing ICU patient data responsibly under the Society of Critical Care Medicine/European Society of Intensive Care Medicine Joint Data Science Collaboration: The Amsterdam University Medical Centers Database (AmsterdamUMCdb) example. Critical Care Medicine, 49(6), e563–e577.
Torio, C. M., & Moore, B. J. (2016). National inpatient hospital costs: The most expensive conditions by payer, 2013 (HCUP Statistical Brief No. 204). Agency for Healthcare Research and Quality.
U.S. Congress. (2016). 21st Century Cures Act, Pub. L. No. 114-255.
U.S. Food and Drug Administration. (2021). Artificial intelligence/machine learning (AI/ML)-based software as a medical device (SaMD) action plan.
U.S. Food and Drug Administration. (2023). Marketing submission recommendations for a predetermined change control plan for artificial intelligence/machine learning-enabled device software functions: Draft guidance for industry and Food and Drug Administration staff.
U.S. Food and Drug Administration. (2024). De Novo classification request for Sepsis ImmunoScore (DEN230036).
U.S. Food and Drug Administration. (2025). 510(k) clearance for TriVerity (K241676).
Vaid, A., Jaladanki, S. K., Xu, J., et al. (2021). Federated learning of electronic health records to improve mortality prediction in hospitalized patients with COVID-19: Machine learning approach. JMIR Medical Informatics, 9(1), Article e24207.
van de Sande, D., van Genderen, M. E., Smit, J. M., et al. (2022). Developing, implementing and governing artificial intelligence in medicine: A step-by-step approach to prevent an artificial intelligence winter. BMJ Health & Care Informatics, 29(1), Article e100495.
Vasey, B., Nagendran, M., Campbell, B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine, 28, 924–933.
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
Winslow, C. J., Edelson, D. P., Churpek, M. M., et al. (2022). The impact of a machine learning early warning score on hospital mortality: A multicenter clinical intervention trial. Critical Care Medicine, 50(9), 1339–1347.
Wolff, R. F., Moons, K. G. M., Riley, R. D., et al. (2019). PROBAST: A tool to assess the risk of bias and applicability of prediction model studies. Annals of Internal Medicine, 170(1), 51–58.
Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065–1070.
Zielinski, C., Winker, M. A., Aggarwal, R., et al. (2023). Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intelligence in relation to scholarly publications. World Association of Medical Editors.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Michal Parzniewski, Natalia Ostruszka, Maria Marusińska, Michał Karpinski, Wiktoria Czyż, Sabina Kolawa, Wojciech Markiewicz, Arnold Borowiec, Julia Pilecka, Jędrzej Wojciechowski

This work is licensed under a Creative Commons Attribution 4.0 International License.
All articles are published in open-access and licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). Hence, authors retain copyright to the content of the articles.
CC BY 4.0 License allows content to be copied, adapted, displayed, distributed, re-published or otherwise re-used for any purpose including for adaptation and commercial use provided the content is attributed.

