PH-AIDE: A Multi-Modal Explainable AI Framework for ICU Risk Stratification and Evidence-Based Public Health Policy Design

Authors

  • Safkat Mustakim Nirjash Author https://orcid.org/0009-0002-4637-4924
    • Formal Analysis
    • Investigation
    • Methodology
    • Writing – Original Draft Preparation
    • Writing – Review & Editing
    • Visualization
  • Sharmin Akter Author
    • Conceptualization
    • Methodology
    • Resources
    • Writing – Original Draft Preparation
    • Writing – Review & Editing
  • Shakim Ahamed Author
    • Dr Emon Hasan Phd Student Author

      DOI:

      https://doi.org/10.70715/jitcai.2026.v3.i4.087

      Keywords:

      Multi-modal AI, ICU risk stratification, Explainable AI, XGBoost, Bi-LSTM, BioBERT, Attention fusion, SHAP, Fairness-aware evaluation, Public health policy, Ablation study

      Abstract

      We present PH-AIDE (Public Health AI Decision Engine), a multi-modal, explainable AI framework for five-class intensive care unit (ICU) risk stratification and evidence-based public health policy design. Drawing exclusively on MIMIC-III critical care data, PH-AIDE fuses three complementary modalities from the same patient cohort: (i) structured clinical features (vitals, laboratory values, ICD-9 diagnoses, comorbidity indices) processed by an XGBoost classifier; (ii) hourly vital-sign time-series from the first 48 hours of ICU admission processed by a two-layer bidirectional LSTM (Bi-LSTM) network with additive attention; and (iii) temporally filtered clinical progress and nursing notes (timestamped within the first 48 hours of admission, with discharge summaries explicitly excluded to prevent information leakage) processed by a fine-tuned BioBERT encoder. Module outputs are fused by a two-layer attention-weighted gating network trained in a staged procedure. Evaluated under temporally blocked cross-validation stratified by risk class, PH-AIDE achieves a macro-averaged F1 score of 0.88 ± 0.01, AUC-ROC of 0.93 ± 0.01, and Brier score of 0.081 ± 0.01 across five risk strata, statistically significantly outperforming all single-modality and naïve ensemble baselines (p < 0.01, paired bootstrap, B = 10,000). A comprehensive ablation study confirms the independent contribution of each modality and the superiority of learned attention weighting over mean fusion. On a separate CDC FluSight walk-forward validation, a standalone Bi-LSTM achieves a 1-week-ahead MAE of 0.38% versus 0.51% for ARIMA and 0.41% for the FluSight-ensemble benchmark. An integrated policy simulation layer using constrained linear programming translates probabilistic risk forecasts into SDG-3-aligned resource allocation recommendations, with WHO Global Health Observatory indicators providing cross-national context. Dual-granularity SHAP explainability (global beeswarm + patient-level waterfall) and subgroup-stratified fairness evaluation with multi-class demographic parity difference (macro-DPD ≤ 0.04) are reported. An extended discussion addresses ethical considerations, governance alignment with WHO AI principles, and reproducibility.

      Downloads

      Download data is not yet available.

      References

      [1] Olawade, D.B., et al. "Using artificial intelligence to improve public health: a narrative review." Frontiers in Public Health 11, art. 1196397 (2023). https://doi.org/10.3389/fpubh.2023.1196397 DOI: https://doi.org/10.3389/fpubh.2023.1196397

      [2] Dixon, S.R., et al. "Unveiling the influence of AI predictive analytics on patient outcomes." Cureus 16(5), e59954 (2024). https://doi.org/10.7759/cureus.59954 DOI: https://doi.org/10.7759/cureus.59954

      [3] Calli, E., et al. "Artificial intelligence in predictive healthcare: a systematic review." Journal of Clinical Medicine 14(19), 6752 (2025). https://doi.org/10.3390/jcm14196752 DOI: https://doi.org/10.3390/jcm14196752

      [4] Flahault, A., et al. "Artificial intelligence in public health: promises, challenges, and an agenda." The Lancet Public Health 10(3), e208–e216 (2025). https://doi.org/10.1016/S2468-2667(24)00298-X

      [5] Nasir, A., et al. "Advancement in public health through machine learning." Journal of Big Data 12, art. 78 (2025). https://doi.org/10.1186/s40537-025-01106-4

      [6] Sinha, S., et al. "Integrating artificial intelligence with mechanistic epidemiological modeling." Nature Communications 16, art. 597 (2025). https://doi.org/10.1038/s41467-024-55634-y DOI: https://doi.org/10.1038/s41467-024-55461-x

      [7] Harishbhai Tilala, M., et al. "Ethical considerations in the use of artificial intelligence in health care." Cureus 16(6), e62443 (2024). https://doi.org/10.7759/cureus.62443 DOI: https://doi.org/10.7759/cureus.62443

      [8] Mushtaq, A., et al. "Personalized health monitoring using explainable AI." Scientific Reports 15, art. 29835 (2025). https://doi.org/10.1038/s41598-025-85959-4

      [9] Dankwa-Mullan, I. "Health equity and ethical considerations in using AI in public health." Preventing Chronic Disease 21, art. 240245 (2024). https://doi.org/10.5888/pcd21.240245 DOI: https://doi.org/10.5888/pcd21.240245

      [10] Wiens, J., et al. "Diagnosing bias in data-driven algorithms for healthcare." Nature Medicine 26(1), 25–26 (2020). https://doi.org/10.1038/s41591-019-0726-6 DOI: https://doi.org/10.1038/s41591-019-0726-6

      [11] Centers for Disease Control and Prevention. "CDC’s vision for using artificial intelligence in public health." Data Modernization Initiative (2025). https://www.cdc.gov/data-strategy/php/ai/index.html

      [12] Rajkomar, A., Dean, J., and Kohane, I. "Machine learning in medicine." New England Journal of Medicine 380(14), 1347–1358 (2019). https://doi.org/10.1056/NEJMra1814259 DOI: https://doi.org/10.1056/NEJMra1814259

      [13] Chen, T. and Guestrin, C. "XGBoost: A scalable tree boosting system." Proc. 22nd ACM SIGKDD, pp. 785–794 (2016). https://doi.org/10.1145/2939672.2939785 DOI: https://doi.org/10.1145/2939672.2939785

      [14] Chae, S., et al. "Predicting infectious disease using deep learning and big data." International Journal of Environmental Research and Public Health 15(8), art. 1596 (2018). https://doi.org/10.3390/ijerph15081596 DOI: https://doi.org/10.3390/ijerph15081596

      [15] Centers for Disease Control and Prevention. "National Syndromic Surveillance Program (NSSP)." https://www.cdc.gov/nssp/index.html (accessed Jan. 2025).

      [16] Shi, Y., et al. "Three-month real-time dengue forecast models." Environmental Health Perspectives 124(9), 1369–1375 (2016). https://doi.org/10.1289/ehp.1509981 DOI: https://doi.org/10.1289/ehp.1509981

      [17] Amann, J., et al. "Explainability for artificial intelligence in healthcare." BMC Medical Informatics and Decision Making 20, art. 310 (2020). https://doi.org/10.1186/s12911-020-01332-6 DOI: https://doi.org/10.1186/s12911-020-01332-6

      [18] Lundberg, S.M. and Lee, S.-I. "A unified approach to interpreting model predictions." Advances in Neural Information Processing Systems (NeurIPS) 30, 4765–4774 (2017).

      [19] Mehrabi, N., et al. "A survey on bias and fairness in machine learning." ACM Computing Surveys 54(6), art. 115 (2021). https://doi.org/10.1145/3457607 DOI: https://doi.org/10.1145/3457607

      [20] World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance. Geneva: WHO (2021).

      [21] World Health Organization. Ethics and governance of AI for health: guidance on large multi-modal models. Geneva: WHO (2024).

      [22] Char, D.S., et al. "Identifying ethical considerations for machine learning healthcare applications." American Journal of Bioethics 20(11), 7–17 (2020). https://doi.org/10.1080/15265161.2020.1819469 DOI: https://doi.org/10.1080/15265161.2020.1819469

      [23] Johnson, A.E.W., et al. "MIMIC-III, a freely accessible critical care database." Scientific Data 3, art. 160035 (2016). https://doi.org/10.1038/sdata.2016.35 DOI: https://doi.org/10.1038/sdata.2016.35

      [24] Centers for Disease Control and Prevention. "FluSight: Flu Forecasting." https://www.cdc.gov/flu/weekly/flusight/index.html (accessed Jan. 2025).

      [25] World Health Organization. "Global Health Observatory (GHO) data." https://www.who.int/data/gho (accessed Jan. 2025).

      [26] van Buuren, S. and Groothuis-Oudshoorn, K. "mice: Multivariate imputation by chained equations in R." Journal of Statistical Software 45(3), 1–67 (2011). https://doi.org/10.18637/jss.v045.i03 DOI: https://doi.org/10.18637/jss.v045.i03

      [27] Guo, C. and Berkhahn, F. "Entity embeddings of categorical variables." arXiv:1604.06737 (2016).

      [28] Charlson, M.E., et al. "A new method of classifying prognostic comorbidity in longitudinal studies." Journal of Chronic Diseases 40(5), 373–383 (1987). https://doi.org/10.1016/0021-9681(87)90171-8 DOI: https://doi.org/10.1016/0021-9681(87)90171-8

      [29] Cleveland, R.B., et al. "STL: A seasonal-trend decomposition procedure based on loess." Journal of Official Statistics 6(1), 3–73 (1990).

      [30] Lee, J., et al. "BioBERT: a pre-trained biomedical language representation model." Bioinformatics 36(4), 1234–1240 (2020). https://doi.org/10.1093/bioinformatics/btz682 DOI: https://doi.org/10.1093/bioinformatics/btz682

      [31] Bergstra, J., Yamins, D., and Cox, D.D. "Making a science of model search." Proc. 30th ICML, pp. 115–123 (2013).

      [32] Hochreiter, S. and Schmidhuber, J. "Long short-term memory." Neural Computation 9(8), 1735–1780 (1997). https://doi.org/10.1162/neco.1997.9.8.1735 DOI: https://doi.org/10.1162/neco.1997.9.8.1735

      [33] Howard, J. and Ruder, S. "Universal language model fine-tuning for text classification." Proc. ACL, pp. 328–339 (2018). https://doi.org/10.18653/v1/P18-1031 DOI: https://doi.org/10.18653/v1/P18-1031

      [34] Baker, R.E., et al. "Mechanistic models versus machine learning." Biology Letters 14(5), art. 20170660 (2018). https://doi.org/10.1098/rsbl.2017.0660 DOI: https://doi.org/10.1098/rsbl.2017.0660

      [35] Khoury, M.J., et al. "Health equity in the implementation of genomics and precision medicine." Genetics in Medicine 24(8), 1630–1639 (2022). https://doi.org/10.1016/j.gim.2022.04.009 DOI: https://doi.org/10.1016/j.gim.2022.04.009

      [36] United Nations. "Sustainable Development Goal 3." https://sdgs.un.org/goals/goal3 (accessed Jan. 2025).

      [37] Centers for Disease Control and Prevention. "Public Health Data Strategy, 2024–2025." https://www.cdc.gov/ophdst/public-health-data-strategy/index.html (accessed Jan. 2025).

      [38] Abadi, M., et al. "Deep learning with differential privacy." Proc. ACM CCS, pp. 308–318 (2016). https://doi.org/10.1145/2976749.2978318 DOI: https://doi.org/10.1145/2976749.2978318

      [39] Pearl, J. Causality: Models, Reasoning, and Inference, 2nd ed. Cambridge University Press (2009). DOI: https://doi.org/10.1017/CBO9780511803161

      [40] Hinton, G., Vinyals, O., and Dean, J. "Distilling the knowledge in a neural network." arXiv:1503.02531 (2015).

      [41] European Commission. "The Artificial Intelligence Act." https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (accessed Jan. 2025).

      [42] Deb, K., et al. "A fast and elitist multiobjective genetic algorithm: NSGA-II." IEEE Transactions on Evolutionary Computation 6(2), 182–197 (2002). https://doi.org/10.1109/4235.996017 DOI: https://doi.org/10.1109/4235.996017

      [43] Daughton, A.R., et al. "An internet-based participatory epidemiology study of influenza." BMC Public Health 17, art. 929 (2017). https://doi.org/10.1186/s12889-017-4896-6

      Downloads

      Published

      07/31/2026

      Data Availability Statement

      All data available upon request.

      How to Cite

      [1]
      Safkat Mustakim Nirjash, Sharmin Akter, S. Ahamed, and D. E. Hasan, “PH-AIDE: A Multi-Modal Explainable AI Framework for ICU Risk Stratification and Evidence-Based Public Health Policy Design”, Journal of Information Technology, Cybersecurity, and Artificial Intelligence, vol. 3, no. 4, pp. 48–66, Jul. 2026, doi: 10.70715/jitcai.2026.v3.i4.087.

      Share